Principal Infrastructure Engineer, AI Cluster Performance & Validation
AI summary of the role
Principal-level role owning performance validation and acceptance of multi-thousand-GPU AI/HPC clusters at Nscale.
What you’ll do
- Define acceptance criteria and performance bars (collective bandwidth, job goodput, model FLOPs utilization) for production-ready clusters.
- Run and instrument real distributed training/inference jobs across thousands of accelerators to validate cluster behavior.
- Lead deep diagnosis of cluster failures and performance regressions across GPU, fabric, storage, and scheduler stack.
- Design and automate validation/burn-in systems (NCCL/RCCL sweeps, HPL/HPCG, MLPerf-style benchmarks, thermal/power soaks).
What you’ll bring
- 10+ years building/operating/debugging large-scale compute infrastructure, with staff/principal-level cross-team technical direction.
- Hands-on experience running real AI compute jobs at scale (pre-training, fine-tuning, or large-scale inference) with distributed training frameworks like PyTorch, Megatron-LM, DeepSpeed.
- Demonstrated experience validating and accepting clusters of thousands of GPUs for performance and reliability.
- Deep understanding of high-performance fabrics (InfiniBand/RoCEv2, RDMA, GPUDirect, adaptive routing, congestion control).
Technologies
NCCL · RCCL · InfiniBand · RoCEv2 · GPUDirect RDMA · PyTorch · Megatron-LM · DeepSpeed · SLURM · Kubernetes · Nsight Systems · Nsight Compute
Source and classification
Internal deployment & tooling · Evidence for this classification:
Overview As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs. Key Responsibilities Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job
More from the job description
Overview As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs. Key Responsibilities Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership. Run and instrument real AI workloads as a diagnostic instrument. Stand up and execute distributed training and inference jobs — open-source and customer-representative models — across thousands of accelerators to validate cluster behavior under genuine load rather than synthetic pro [... source excerpt omitted ...] ettings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput. Partner with Infrastructure, Platform, SRE, and customer-facing teams to translate operational and customer performance needs into durable engineering solutions, and to feed diagnostic signal back into provisioning, remediation, and capacity workflows. Build production-grade Python systems and performance tooling for automated triage, telemetry correlation, and regression detection, leveraging AI tools to accelerate delivery. Assess impact to the team's software and validation stack from new hardware product programs, and explore AI-driven process i [... source excerpt omitted ...] tional excellence across all services, and raise the diagnostic capability of the wider organization through mentorship, runbooks, and post-incident technical write-ups. Required Qualifications Education: Bachelor's or higher degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Experience: 10+ years of relevant experience building, operating, or debugging large-scale compute infrastructure, including significant time at staff or principal level owning cross-team technical direction. AI Workload Expertise: Hands-on experience running real AI compute jobs at scale — pre-training, fine-tuning, or large-scale inference of open-
Employer postings · Data from · Sources