Senior ML Infrastructure Engineer (Compute)
AI summary of the role
Senior ML Infrastructure Engineer on the AI Validation Platform team, building and scaling compute infrastructure for autonomous vehicle simulation (L3/L4/L5).
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
What you’ll do
- Design and implement core platform backend software components.
- Collaborate with Simulation engineers, ML engineers and researchers to understand critical workflows and deliver incremental value.
- Lead technical decision-making on Compute architecture, cloud capacity provisioning, caching, and auto-scaling mechanisms.
- Drive development of monitoring, observability, and metrics for reliability and resource optimization.
What you’ll bring
- 4+ years of industry experience with high performance backend services.
- Strong expertise in Go or similar coding languages.
- Experience with cloud platforms such as GCP, Azure, or AWS.
- Experience delivering cross-functional initiatives.
Technologies
Go · GCP · Azure · AWS · Google Compute Engine · GPUs · HPC · distributed systems · cloud capacity provisioning · telemetry
About General Motors
Global automaker selling Chevrolet, GMC, Cadillac, and Buick vehicles, plus financing and connected services, while shifting into EVs and driver assistance.
Public · 5000+ people
Source and classification
Internal deployment & tooling · Evidence for this classification:
Job Description About the Team: The AI Validation Platform team owns the cloud-agnostic, reliable, and cost-efficient platform that powers GM’s AV efforts. We’re proud to serve as the infrastructure platform for teams developing autonomous vehicles (L3/L4/L5). Our platform supports the simulated validation of state-of-the-art (SOTA) machine learning models, with a focus on performance, availability, concurrency, and scalability. We enable rapid innovation and development by prioritizing high-impact, ML-centric use cases. About the Role: We are seeking a Senior ML Infrastructure engineer to help build and scale robust Compute platforms for Simulation workflows. In this role, you will focus on scaling, driving efficiency, and high utilization of cutting-edge GPUs, while also leveling up the platform’s reliability. The successful candidate will have experience building and running
More from the job description
Job Description About the Team: The AI Validation Platform team owns the cloud-agnostic, reliable, and cost-efficient platform that powers GM’s AV efforts. We’re proud to serve as the infrastructure platform for teams developing autonomous vehicles (L3/L4/L5). Our platform supports the simulated validation of state-of-the-art (SOTA) machine learning models, with a focus on performance, availability, concurrency, and scalability. We enable rapid innovation and development by prioritizing high-impact, ML-centric use cases. About the Role: We are seeking a Senior ML Infrastructure engineer to help build and scale robust Compute platforms for Simulation workflows. In this role, you will focus on scaling, driving efficiency, and high utilization of cutting-edge GPUs, while also leveling up the platform’s reliability. The successful candidate will have experience building and running scalable distributed systems. They will rapidly test and promote ideas, have strong problem-solving skills, and demonstrate a bias for action. You will play a key role in shaping the architecture, roadmap, and user experience of a robust service supporting our AI Validation / Simulation needs. The ideal candidate brings experience in designing distributed systems, strong problem-solving skills, and a get-it-done attitude. This is a high-impact opportunity to influence the future of AI infrastructure [... source excerpt omitted ...] implement core platform backend software components. Collaborate with Simulation engineers, ML engineers and researchers to understand critical workflows, parse them to platform requirements, and deliver incremental value. Lead technical decision-making on Compute architecture, cloud capacity provisioning, caching, and auto-scaling mechanisms. Drive the development of monitoring, observability, and metrics to ensure reliability, performance, and resource optimization. Proactively research and integrate frameworks, hardware accelerators, and distributed computing techniques. Lead large-scale technical initiatives across GM’s ML infrastructure. Raise the engineering bar through [... source excerpt omitted ...] e Compute Engine. Experience with hardware-in-the-loop validation systems. Experience with high performance computing (HPC). Experience working with or designing interfaces and clients for developer workflows. Familiarity with telemetry, and other feedback loops to inform product improvements. Familiarity with hardware acceleration (GPUs) and optimizations. Why Join Us? If you’re excited to tackle some of today’s most complex engineering challenges, see the impact of your work in real-world AV applications, and help shape the future of AI infrastructure at GM—this is the team for you. About GM Our vision is a world with Zero Crashes, Zero Emissions and Zero Congestion
Employer postings · Data from · Sources