Skip to main content
Loading
loadingbar
Loading, Please wait..!!

Technical Program Manager- AI Cluster Validation

  • Job type Posted on: Jul 19, 2026
  • Experience level AMD
  • Employment type Austin, Texas
  • Employment type Onsite
  • Salary Full-time

Point Apply Here APPLY LATER

Curious about compensation?

Explore the historical salary trends, average pay, and estimated compensation for Technical Program Manager- AI Cluster Validation roles in Texas.

View Salary Guide →

Job Title :

Technical Program Manager- AI Cluster Validation

Job Type :

Full-time

Job Location :

Austin Texas United States

Remote :

No

Jobcon Logo Job Description :

Technical Program Manager – AI Cluster Validation We are seeking a Technical Program Manager to lead execution of AI cluster engineering programs with deep focus on GPU platforms, rack-level solutions, and AI Cluster validation. This role is responsible for driving end-to-end delivery from GPU + server integration through rack bring‑up, scale testing, failure analysis, and system debug closure, ensuring platform readiness for hyperscale and enterprise AI deployments. Overview At AMD, our mission is to build great products that accelerate next‑generation computing experiences—from AI and data centers to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. Our culture pushes the limits of innovation to solve the world’s most important challenges, striving for execution excellence and inclusivity. The Person You are a hands‑on TPM who thrives in complex, fast‑moving ecosystems and can connect deep technical details to crisp program plans, executive reporting and customer outcomes. You are comfortable driving execution in bring‑up and EVT/DVT/PVT, unblocking debug, and making data‑driven trade‑offs to keep programs moving. You bring urgency, ownership and clarity to ambiguous problem spaces and communicate effectively from lab floor to executive review. Key Responsibilities Program Leadership & Execution Define, plan and drive program plans for AI infrastructure systems validation and readiness, including server integration, rack bring‑up and cluster‑scale deployment readiness. Create and maintain core PM artifacts: schedules, dependency maps, resource forecasts, risk/issue logs and program dashboards/status reports. Identify and drive mitigation plans for issues/risks, including cross‑team escalations and corrective actions across multiple engineering areas. Drive regular execution reviews with engineering teams and provide concise, data‑driven updates to senior leadership. GPU & Platform Execution Own program execution for GPU‑based AI platforms, spanning system bring‑up, qualification, scale readiness and deployment validation across server, rack and cluster levels. Drive alignment across GPU, CPU, firmware, BIOS/BMC and system teams to ensure readiness for scale testing and customer workloads. Track platform issues, and debug dependencies; ensure risks are clearly documented, owned and mitigated. Coordinate inter‑disciplinary teams to unblock rack access, power, networking and test readiness. AI Rack / Cluster Validation Own program planning and execution for multi‑node and multi‑rack scale testing, including test strategy, scheduling, coverage tracking and readiness gates. Lead end‑to‑end delivery of rack‑level AI solutions, including compute trays, switch trays, cabling, power, cooling and management infrastructure. Ensure rack bring‑up plans are executable, resourced and gated with clear entry/exit criteria across EVT, DVT and scale phases. Partner with scale, performance and automation teams to ensure workloads, stress tests and regression plans are ready before hardware arrives. Debug, Failure Analysis & Risk Management Lead platform debug, coordinating across engineering teams to ensure fast triage, root‑cause analysis and resolution of system‑level issues. Track high‑impact failures (GPU, HSIO, FW, rack, network) through debug forums ensuring clear ownership and closure plans. Balance debug depth versus program timelines, escalating trade‑offs when needed to keep leadership informed of risk and impact. Required Qualifications Experience leading complex hardware or AI infrastructure programs with ownership of bring‑up, validation and deployment phases. Strong technical understanding of GPU‑based AI systems, rack architectures and datacenter infrastructure. Proven ability to manage ambiguity, drive debug execution and lead cross‑functional teams without direct authority. Strong written and verbal communication skills, including executive‑level status reporting. Proficiency with program management and execution tools (Jira, Confluence, dashboards, Excel/PowerPoint). Preferred Qualifications Hands‑on experience with GPU cluster scale testing, system stress or performance validation. Familiarity with rack‑level bring‑up, power/cooling constraints, networking and failure modes at scale. Experience working through hardware/firmware debug cycles in pre‑production or customer‑facing environments. Academic Credentials Bachelor’s or master’s degree in systems, EE, CS or related engineering discipline. PMP, Scrum Master or equivalent program management training. Location Austin, TX Benefits Benefits offered are described: AMD benefits at a glance. Equal Opportunity Statement AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. This posting is for an existing vacancy. This role is not eligible for visa sponsorship. #J-18808-Ljbffr

Jobcon Logo Position Details

Posted:

Jul 19, 2026

Reference Number:

14660_B15FD6CB1EBA08FBAD2CE96FC861F3FD

Employment:

Full-time

Salary:

Not Available

City:

Austin

Job Origin:

APPCAST_CPC

Share this job:

  • linkedin

Jobcon Logo
A job sourcing event
In Dallas Fort Worth
Aug 19, 2017 9am-6pm
All job seekers welcome!

Technical Program Manager- AI Cluster Validation    Apply

Click on the below icons to share this job to Linkedin, Twitter!

Technical Program Manager – AI Cluster Validation We are seeking a Technical Program Manager to lead execution of AI cluster engineering programs with deep focus on GPU platforms, rack-level solutions, and AI Cluster validation. This role is responsible for driving end-to-end delivery from GPU + server integration through rack bring‑up, scale testing, failure analysis, and system debug closure, ensuring platform readiness for hyperscale and enterprise AI deployments. Overview At AMD, our mission is to build great products that accelerate next‑generation computing experiences—from AI and data centers to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. Our culture pushes the limits of innovation to solve the world’s most important challenges, striving for execution excellence and inclusivity. The Person You are a hands‑on TPM who thrives in complex, fast‑moving ecosystems and can connect deep technical details to crisp program plans, executive reporting and customer outcomes. You are comfortable driving execution in bring‑up and EVT/DVT/PVT, unblocking debug, and making data‑driven trade‑offs to keep programs moving. You bring urgency, ownership and clarity to ambiguous problem spaces and communicate effectively from lab floor to executive review. Key Responsibilities Program Leadership & Execution Define, plan and drive program plans for AI infrastructure systems validation and readiness, including server integration, rack bring‑up and cluster‑scale deployment readiness. Create and maintain core PM artifacts: schedules, dependency maps, resource forecasts, risk/issue logs and program dashboards/status reports. Identify and drive mitigation plans for issues/risks, including cross‑team escalations and corrective actions across multiple engineering areas. Drive regular execution reviews with engineering teams and provide concise, data‑driven updates to senior leadership. GPU & Platform Execution Own program execution for GPU‑based AI platforms, spanning system bring‑up, qualification, scale readiness and deployment validation across server, rack and cluster levels. Drive alignment across GPU, CPU, firmware, BIOS/BMC and system teams to ensure readiness for scale testing and customer workloads. Track platform issues, and debug dependencies; ensure risks are clearly documented, owned and mitigated. Coordinate inter‑disciplinary teams to unblock rack access, power, networking and test readiness. AI Rack / Cluster Validation Own program planning and execution for multi‑node and multi‑rack scale testing, including test strategy, scheduling, coverage tracking and readiness gates. Lead end‑to‑end delivery of rack‑level AI solutions, including compute trays, switch trays, cabling, power, cooling and management infrastructure. Ensure rack bring‑up plans are executable, resourced and gated with clear entry/exit criteria across EVT, DVT and scale phases. Partner with scale, performance and automation teams to ensure workloads, stress tests and regression plans are ready before hardware arrives. Debug, Failure Analysis & Risk Management Lead platform debug, coordinating across engineering teams to ensure fast triage, root‑cause analysis and resolution of system‑level issues. Track high‑impact failures (GPU, HSIO, FW, rack, network) through debug forums ensuring clear ownership and closure plans. Balance debug depth versus program timelines, escalating trade‑offs when needed to keep leadership informed of risk and impact. Required Qualifications Experience leading complex hardware or AI infrastructure programs with ownership of bring‑up, validation and deployment phases. Strong technical understanding of GPU‑based AI systems, rack architectures and datacenter infrastructure. Proven ability to manage ambiguity, drive debug execution and lead cross‑functional teams without direct authority. Strong written and verbal communication skills, including executive‑level status reporting. Proficiency with program management and execution tools (Jira, Confluence, dashboards, Excel/PowerPoint). Preferred Qualifications Hands‑on experience with GPU cluster scale testing, system stress or performance validation. Familiarity with rack‑level bring‑up, power/cooling constraints, networking and failure modes at scale. Experience working through hardware/firmware debug cycles in pre‑production or customer‑facing environments. Academic Credentials Bachelor’s or master’s degree in systems, EE, CS or related engineering discipline. PMP, Scrum Master or equivalent program management training. Location Austin, TX Benefits Benefits offered are described: AMD benefits at a glance. Equal Opportunity Statement AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. This posting is for an existing vacancy. This role is not eligible for visa sponsorship. #J-18808-Ljbffr

Loading
Please wait..!!