Skip to content
NEXTMOVEFDE careers · United States

Cloud & Customer Solutions Engineer - DC GPU

AI summary of the role

Embedded production engineer on AMD's Applied AI team, owning end-to-end deployment and operation of AMD Instinct GPU clusters at strategic AI customers (frontier labs, NeoCloud, CSPs).

What you’ll do

  • Own customer deployments end-to-end: cluster bring-up, burn-in, production readiness certification, workload onboarding, performance validation, and sustained operation on AMD Instinct GPU fleets
  • Deploy and tune large-scale training/inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer SLOs across cloud, NeoCloud, and bare-metal
  • Lead root-cause analysis and resolution of production incidents, including Sev-1 response, driving fixes to permanent closure
  • Deploy agentic AI solutions into customer environments and own their production behavior

What you’ll bring

  • 5+ years of production software or infrastructure engineering, including operating/deploying systems in environments you did not build
  • Hands-on GPU compute at scale: cluster deployment, distributed training or high-throughput inference, performance debugging, workload optimization
  • Strong knowledge of Kubernetes/Slurm, containerized GPU workloads, RCCL/NCCL, RoCE/InfiniBand, Prometheus/Grafana
  • Cloud platform depth (AWS, Azure, GCP, or NeoCloud) including hybrid and bare-metal patterns

Technologies

ROCm · AMD Instinct · vLLM · SGLang · RCCL · NCCL · Kubernetes · Slurm · RoCE · InfiniBand · Prometheus · Grafana

Source and classification

Production engineering · Evidence for this classification:

with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people. THE ROLE: As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers — frontier labs, NeoCloud providers, CSPs, and AI-native companies — to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets. To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer production
How jobs are selected

Employer postings · Data from · Sources