CVP of Applied AI FDE
AI summary of the role
This is a senior executive role (CVP) leading AMD's Field Deployment Engineering (FDE) organization, responsible for building and scaling a team that covers the full customer deployment lifecycle for massive GPU clusters.
What you’ll do
- Build and scale a world-class FDE organization combining ML Generalists, Low-Level Kernel Optimizers, and Solutions Architects.
- Define and institutionalize the FDE Engagement Model to maximize resource leverage and ensure consistent, high-velocity customer outcomes.
- Oversee technical onboarding of massive GPU clusters, troubleshooting collective communication errors and optimizing training/inference strategies.
- Drive and maintain industry-leading Customer GPU Utilization across clusters of thousands of GPUs.
What you’ll bring
- Demonstrated track record leading high-impact technical teams in high-stakes environments (Cloud Infrastructure, AI Platform, or HPC).
- Deep understanding of the hardware/software stack from the metal up.
- Commercial acumen with understanding of ARR, Churn, Margin and impact on deal velocity.
- Experience leading through Sev0 customer incidents with executive communication and rapid root cause resolution.
Technologies
PyTorch · JAX · TensorFlow · Slurm · Ray · Kubernetes · Docker · CUDA · Nsight Systems · Triton Inference Server · vLLM · TensorRT-LLM
Source and classification
Production engineering · Evidence for this classification:
Optimizers, and Solutions Architects to cover the full customer deployment lifecycle. This is a highly visible role with large scope and impact. THE PERSON: Define and institutionalize the FDE Engagement Model to maximize resource leverage and ensure consistent, high-velocity customer outcomes. Serve as the Voice of the Customer internally: Translate field intelligence and customer challenges into concrete, prioritized engineering roadmaps, and ensure execution. KEY RESPONSIBILITIES: Cluster Bring-up & Optimization: Oversee the technical onboarding of massive GPU clusters. Ensure your team can troubleshoot collective communication errors, debug framework issues, and optimize training/inference strategies. Utilization Engineering (The North Star Metric): Drive and maintain industry-leading Customer GPU Utilization across clusters of thousands of GPUs, making cluster satisfaction the
More from the job description
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. This role is not eligible for visa sponsorship. THE ROLE: Build and scale a world-class FDE organization, strategically combining ML Generalists, Low-Level Kernel Optimizers, and Solutions Architects to cover the full customer deployment lifecycle. This is a highly visible role with large scope and impact. THE PERSON: Define and institutionalize the FDE Engagement Model to maximize resource leverage and ensure consistent, high-velocity customer outcomes. Serve as the Voice of the Customer internally: Translate field intelligence and customer challenges into concrete, prioritized engineering roadmaps, and ensure execution. KEY RESPONSIBILITIES: Cluster [... source excerpt omitted ...] ive communication errors, debug framework issues, and optimize training/inference strategies. Utilization Engineering (The North Star Metric): Drive and maintain industry-leading Customer GPU Utilization across clusters of thousands of GPUs, making cluster satisfaction the key measure of success. High-Performance Model Deployment: Enable customer success by deeply optimizing open-sourced models (Llama 3, DeepSeek, Mixtral) and proprietary models for our specific hardware topology, utilizing tools like vLLM and TensorRT-LLM. Executive Technical Sponsorship: Act as the technical authority and executive sponsor on large deals, possessing the credibility to validate architecture w [... source excerpt omitted ...] Churn, Margin) and the ability to articulate how technical solutions impact deal velocity and business outcomes. High-Stakes Crisis Management: Experience leading through "Sev0" customer incidents (e.g., massive training run failures), demonstrating the poise and clarity required to manage executive communication while guiding rapid root cause resolution. Technical Competency AI Frameworks: PyTorch, JAX, TensorFlow. Distributed Computing: Slurm, Ray, Kubernetes (K8s), Docker. GPU Ecosystem: NVIDIA drivers, CUDA profiling (Nsight Systems), Triton Inference Server. LLM Operations (Differentiator): Significant experience with advanced LLM deployment and customization techniqu
Employer postings · Data from · Sources