Staff Software Engineer, TPU, Performance
AI summary of the role
Staff-level role in Google's Core ML organization focused on optimizing ML model performance on TPU systems for Gemini and open-source models.
What you’ll do
- Identify and maintain ML training and serving benchmarks representative of Google production and the broader ML industry.
- Achieve performance targets for customer launches and engaged benchmark submissions (ML Commons, InferenceMAX, etc.).
- Use benchmarks to identify performance opportunities and drive out-of-the-box performance improvements in compiler and runtime.
- Engage with product teams and researchers to solve performance problems, including onboarding new models on new TPU hardware and enabling large-scale training on thousands of TPUs.
What you’ll bring
- Bachelor's degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience with ML infrastructure or specialization in an ML field (e.g., speech/audio, reinforcement learning).
- 5 years of experience with ML design and ML infrastructure (model deployment, evaluation, data processing, debugging, fine-tuning).
Technologies
TPU · JAX · PyTorch · MLIR · OpenXLA · Triton · CUDA · OpenCL · GPU · compiler optimization
About Google
Builds global consumer, ads, cloud, developer and AI platforms spanning Search, YouTube, Android, Workspace and Gemini.
Public
Source and classification
Internal deployment & tooling · Evidence for this classification:
full-stack as we continue to push technology forward. Google’s Core Machine Learning (ML) organization is seeking software engineers to join the team known for pioneering work with Tensor Processing Units (TPUs). In this role, you will work on Gemini, as well as industry leading open-source models, to understand model architecture and optimize the performance of these Machine Learning (ML) models on TPU systems for both Just After eXecution (JAX) and PyTorch platforms. You will improve the performance of ever-evolving ML workloads, achieving results. These fundamental efforts will influence next-generation (next-gen) TPU architectures via partnerships, ensuring performance for Gemini and Open-Source Software (OSS) Machine Learning (ML) models.The Core team builds the technical foundation behind Google’s flagship products. We are owners and advocates for the underlying design elements,
More from the job description
About the job Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward. Google’s Core Machine Learning (ML) organization is seeking software engineers to join the team known for pioneering work with Tensor Processing Units (TPUs). In this role, you will work on Gemini, as well as industry leading open-source models, to understand model architecture and optimize the performance of these Machine Learning (ML) models on TPU systems for both Just After eXecution (JAX) and PyTorch platforms. You will improve the performance [... source excerpt omitted ...] ding job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits Learn more about benefits at Google. Responsibilities Identify and maintain ML training and serving benchmarks that are representative to Google production and broader ML industry. Achieve performance for customer launches, and in case of third-party/Open-Source Software (3P/OSS) models, for engaged benchmark submissions ML commons, InferenceMAX, etc.). Use the benchmarks to identify performance opportunities and drive out-of-the-box performance toward improving the compiler, runtime, etc., in collaboration with those teams. Engage with Goo [... source excerpt omitted ...] on a very large-scale (i.e., thousands of TPUs). Analyze performance and efficiency metrics to identify bottlenecks, design, and implement solutions at Google fleet-wide scale. Qualifications Minimum qualifications: Bachelor’s degree or equivalent practical experience. 8 years of experience in software development. 5 years of experience with one or more of the following: speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field. 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processin
Employer postings · Data from · Sources