Staff AI/ML Engineer, AI Rapid Response Team
AI summary of the role
Staff-level AI/ML engineer anchoring Google's AI Rapid Response Team, acting as chief efficiency strategist for high-stakes AI workloads.
What you’ll do
- Lead technical architecture for 1-6 month FDE embeds and 2-4 week Strike Sprints on ML efficiency bottlenecks.
- Design and write production C++/Python for model compression, speculative decoding, dynamic batching, and high-throughput serving.
- Diagnose distributed latency/throughput issues across XManager, Pathways, TPU/GPU clusters; implement XLA/Pallas/Custom Op optimizations.
- Translate vague executive mandates into rigorous efficiency scopes with latency, FLOPs, and tokenomics thresholds.
What you’ll bring
- Bachelor's degree or equivalent practical experience.
- 8 years of software development experience.
- 5 years testing/launching software products and 3 years software design/architecture.
- 5 years in speech/audio, reinforcement learning, ML infrastructure, or other ML specialization.
Technologies
C++ · Python · XLA · Pallas · CUDA · Triton · TPU · GPU · XManager · Pathways · FP8 · INT4
About Google
Builds global consumer, ads, cloud, developer and AI platforms spanning Search, YouTube, Android, Workspace and Gemini.
Public
Source and classification
Internal deployment & tooling · Evidence for this classification:
full-stack as we continue to push technology forward. With your technical expertise you will manage project priorities, deadlines, and deliverables. You will design, develop, test, deploy, maintain, and enhance software solutions. On the AI Rapid Response Team, you will serve as the primary technical anchor and chief efficiency strategist for Google's most demanding AI workloads. You will operate as an "ambiguity buster," taking nebulous VP-level mandates around compute constraints, latency spikes, and infrastructure scaling costs, and translating them into crisp, mathematically validated engineering solutions. Trading long-term maintenance of legacy systems for continuous zero-to-one pathfinding velocity, you will lead Strike Sprints and embedded FDE engagements across Google. You will write high-performance production code, architect novel model efficiency pipelines, optimize
More from the job description
About the job Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward. With your technical expertise you will manage project priorities, deadlines, and deliverables. You will design, develop, test, deploy, maintain, and enhance software solutions. On the AI Rapid Response Team, you will serve as the primary technical anchor and chief efficiency strategist for Google's most demanding AI workloads. You will operate as an "ambiguity buster," taking nebulous VP-level mandates around compute constraints, latency spikes, a [... source excerpt omitted ...] and translating them into crisp, mathematically validated engineering solutions. Trading long-term maintenance of legacy systems for continuous zero-to-one pathfinding velocity, you will lead Strike Sprints and embedded FDE engagements across Google. You will write high-performance production code, architect novel model efficiency pipelines, optimize inference serving engines, and establish graceful exit architectures that empower partner teams to run permanently lean. Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help d [... source excerpt omitted ...] ding job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits Learn more about benefits at Google. Responsibilities Lead technical architecture and system design for complex 1–6-month Forward Deployed Engineer (FDE) embeds and 2–4-week Strike Sprints targeting high-leverage Machine Learning efficiency bottlenecks. Design, prototype, and write robust production C++ and Python code for model compression, speculative decoding engines, dynamic batching layers, and high-throughput serving pipelines. Diagnose subtle distributed latency and throughput bottlenecks across XManager, Pathways, and Tensor Processi
Employer postings · Data from · Sources