Skip to content
NEXTMOVEFDE careers · United States

Principal Solutions Engineering – AI server/rack Infrastructure

AI summary of the role

Principal-level technical leadership role in AMD's Data Center Platform Engineering Group, owning system design support, rack-level bring-up, and customer engagement for AMD Instinct AI server/rack solutions.

What you’ll do

  • Architect and optimize rack-scale AI deployments with AMD Instinct GPUs
  • Lead hands-on rack, platform, and component-level debug and validation
  • Drive system firmware (BIOS/BMC) debug and deployment at scale
  • Own end-customer technical support and issue resolution

What you’ll bring

  • Advanced experience in system architecture, hardware/firmware debug, and customer-facing engineering (HPC/AI-ML preferred)
  • Deep understanding of server/rack architecture (x86, GPU, PCIe, interconnects)
  • Strong proficiency in BIOS/UEFI and BMC/OpenBMC debug and deployment
  • Experience with bring-up tools (oscilloscopes, logic analyzers, ITP, JTAG)

Technologies

AMD Instinct · GPU · x86 · PCIe · BIOS/UEFI · BMC/OpenBMC · ITP · JTAG · HPC · AI/ML

Source and classification

Implementation & delivery · Evidence for this classification:

ahead of the next challenge. As experts in engineering, manufacturing, and supply chain, we’re the bridge between problem and solution for the world’s leading OEM & ODM partners and cloud services providers. Our customers depend on us to solve their most complex server/rack design needs. Come and join our Data Center Platform Engineering Group where we are building amazing, powered products with amazing people. THE PERSON: AMD is searching for a dynamic and experienced Principal Member of Technical Staff to own system design support, rack-level bring-up, and critical customer engagement for our cutting-edge AMD Instinct™ product line. In this high-visibility role, you will act as the technical bridge between AMD’s internal system architects, platform development teams, and our OEM partners. You will not only influence the design and architecture of AI solutions but also lead hands-on
More from the job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: AMD's Data Center Platform Engineering Group (DPEG) is designing, developing, and delivering innovative technology infrastructure enabling the digital world. We create cloud-enabling server/rack solutions that help the world’s leading companies turn their ideas into reality. Our customers are future-focused and so are we, always a step ahead of the next challenge. As experts in engineering, manufacturing, and supply chain, we’re the bridge between problem and solution for the world’s leading OEM & ODM partners and cloud services providers. Our customers depend on us to solve their most complex server/rack design needs. Come and join our Data Center Platform Engineering Group where we are building amazing, powered products with amazing people. THE PERSON: AMD is searching for a dynamic and experienced Principal Member of Techni [... source excerpt omitted ...] tinct™ product line. In this high-visibility role, you will act as the technical bridge between AMD’s internal system architects, platform development teams, and our OEM partners. You will not only influence the design and architecture of AI solutions but also lead hands-on debug and validation efforts at customer locations. As a technical leader, you will drive engineering, root cause analysis, and influence future roadmaps based on field execution. KEY RESPONSIBILITIES: System Architecture & Design Support Solution Optimization: Partner deeply with customers to architect and optimize Rack-Scale AI solution deployments using AMD Instinct GPUs. Design Reviews: Provide support [... source excerpt omitted ...] n and deployment of AMD AI platforms. Hands-on Engineering: Drive hands-on rack, platform, and component-level debug and validation. This includes complex stress testing, issue reproductions, and deep-dive root cause analysis. Issue Resolution: Lead customer issue resolution efforts, gathering diagnostics, managing critical escalations, and driving long-term process improvements to ensure customer success. System Firmware Debug & Deployment: Lead debug efforts for system firmware (BIOS, BMC) during initial bring-up and large-scale deployment phases. Ensure seamless integration between hardware, firmware, and software stacks, and resolve interaction issues in customer environment

How jobs are selected

Employer postings · Data from · Sources