Skip to content
NEXTMOVEFDE careers · United States

Senior Solutions Architect - AI Infrastructure

AI summary of the role

NVIDIA seeks a Senior Solutions Architect to design and deploy next-generation GPU clusters for the world's largest AI supercomputers.

What you’ll do

  • Partner with internal engineering on GPU cluster design and networking, conveying architecture to customers and field teams.
  • Guide field teams and customers in cluster design to optimize performance and supportability.
  • Ensure successful first deployments of new products, including new network architectures and topologies.
  • Perform hands-on debugging of cluster design, configuration, and performance issues.

What you’ll bring

  • BS/MS/PhD in CS, EE, CE, Physics, or related field (or equivalent experience).
  • 8+ years experience in cluster design, validation, and issue resolution on GPU and HPC clusters.
  • Proven expertise in designing large-scale distributed systems, AI clusters, or HPC infrastructure.
  • Ability to translate complex engineering concepts into customer-ready documentation.

Technologies

GPU · NVLink · NVIDIA Networking · NCCL · MPI · IMEX · NMX · HPC clusters · distributed training

About NVIDIA

Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.

Public

Source and classification

Implementation & delivery · Evidence for this classification:

designs as well as all of our software solutions directly between engineering and field teams supporting customers with the most demanding requirements. You will work on end-to-end cluster design and architecture, performance modeling, validation, and NPI cluster deployments. Your expertise will directly influence how the world’s leading AI companies, cloud providers, hyperscalers, research institutions, and enterprises build their infrastructure. What you’ll be doing: Partner with internal engineering efforts in GPU cluster design and networking and convey architecture and optimal process information both direct to customer and with field teams supporting customers Guide field teams and their customers in cluster design, weighing design principles but also complex, situational limitations to make the most performant and supportable GPU clusters possible Work closely with field
More from the job description

NVIDIA is building the world’s most groundbreaking and innovative accelerated computing platforms for AI and HPC. Because of our work, scientists, researchers, and engineers can push the boundaries of what’s possible. We pioneered a supercharged form of computing that powers everything from breakthrough AI research to the world’s fastest supercomputers. We are seeking a highly motivated Senior Solutions Architect to join the Cluster Design and Architecture team with a focus on GPU, NVLink, and infrastructure design. In this role, you will be at the forefront of assisting with designs and architectures for some for the largest next-generation GPU-based clusters enabling the world’s most advanced AI supercomputers and enterprise AI infrastructure in the field. As a Solutions Architect, you will serve as a key technical expert bridging NVIDIA’s ground breaking GPU and NVLink technology designs as well as all of our software solutions directly between engineering and field teams supporting customers with the most demanding requirements. You will work on end-to-end cluster design and architecture, performance modeling, validation, and NPI cluster deployments. Your expertise will directly influence how the world’s leading AI companies, cloud providers, hyperscalers, research institutions, and enterprises build their infrastructure. What you’ll be doing: Partner with internal engin [... source excerpt omitted ...] luster design and networking and convey architecture and optimal process information both direct to customer and with field teams supporting customers Guide field teams and their customers in cluster design, weighing design principles but also complex, situational limitations to make the most performant and supportable GPU clusters possible Work closely with field teams supporting customers to ensure successful first deployments with new products, including new network architectures and topologies Feedback customer/field perspectives on cluster design and workflows back to engineering teams designing internal clusters and/or creating customer facing documentation on standard p [... source excerpt omitted ...] ands-on work to assist field teams debugging issues relating to cluster design, configuration, and performance employing internal engineering expertise and known bugs Support NPI customer deployments with new GPU/Networking architectures What we need to see: BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, Physics, or related field (or equivalent experience) 8+ years of experience in cluster design, validation, and issue resolution, specifically on GPU and HPC clusters Proven expertise in designing large-scale distributed systems, AI clusters, or HPC infrastructure Ability to translate sophisticated engineering concepts into customer-ready d

How jobs are selected

Employer postings · Data from · Sources