Cloud & Customer Solutions Engineer - DC GPU
AI summary of the role
Embedded production engineer on AMD's Applied AI team, owning end-to-end deployment and operation of AMD Instinct GPU clusters at strategic AI customers (frontier labs, NeoCloud, CSPs).
What you’ll do
- Own customer deployments end-to-end: cluster bring-up, burn-in, production readiness certification, workload onboarding, performance validation, and sustained operation on AMD Instinct GPU fleets
- Deploy and tune large-scale training/inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer SLOs across cloud, NeoCloud, and bare-metal
- Lead root-cause analysis and resolution of production incidents, including Sev-1 response, driving fixes to permanent closure
- Deploy agentic AI solutions into customer environments and own their production behavior
What you’ll bring
- 5+ years of production software or infrastructure engineering, including operating/deploying systems in environments you did not build
- Hands-on GPU compute at scale: cluster deployment, distributed training or high-throughput inference, performance debugging, workload optimization
- Strong knowledge of Kubernetes/Slurm, containerized GPU workloads, RCCL/NCCL, RoCE/InfiniBand, Prometheus/Grafana
- Cloud platform depth (AWS, Azure, GCP, or NeoCloud) including hybrid and bare-metal patterns
Technologies
ROCm · AMD Instinct · vLLM · SGLang · RCCL · NCCL · Kubernetes · Slurm · RoCE · InfiniBand · Prometheus · Grafana
Source and classification
Production engineering · Evidence for this classification:
with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people. THE ROLE: As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers — frontier labs, NeoCloud providers, CSPs, and AI-native companies — to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets. To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer productionHow jobs are selected
Employer postings · Data from · Sources