Senior Performance Co-Design Engineer, TPU
AI summary of the role
Senior engineer on Google's TPU Chip Architecture and Performance Co-design team, focused on LLM serving performance studies.
What you’ll do
- Conduct comprehensive serving performance studies on current and emerging LLMs (1P/3P models)
- Develop and maintain advanced simulation, profiling, and modeling tools to identify bottlenecks and project serving workload performance
- Partner with model researchers, software and hardware teams to co-design architectural improvements for LLM inference latency and throughput
- Drive data-backed decisions that influence the roadmap for future TPU/Cloud Silicon architectures
What you’ll bring
- Bachelor's degree in CS, EE, CE, or related field, or equivalent practical experience
- 5 years of experience in performance modeling/engineering, computer architecture, co-design, or systems engineering
- Experience programming in C++ or Python
Technologies
TPU · LLM · C++ · Python · simulation · profiling · modeling · inference · serving · accelerator architectures
About Google
Builds global consumer, ads, cloud, developer and AI platforms spanning Search, YouTube, Android, Workspace and Gemini.
Public
Source and classification
Internal deployment & tooling · Evidence for this classification:
About the job Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action. In this role, you aren’t just building technology; you’re shaping the frontier of enterprise and driving the evolution of advanced models. The TPU Chip Architecture and Performance Co-design team is at the forefront of optimizing Google's custom AI silicon for next-generation machine learning models. As a Senior Performance Co-Design Engineer, you will focus and conduct LLM Serving Studies. In this role, you will work on analyzing and optimizing the serving performance of emerging models and
More from the job description
About the job Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action. In this role, you aren’t just building technology; you’re shaping the frontier of enterprise and driving the evolution of advanced models. The TPU Chip Architecture and Performance Co-design team is at the forefront of optimizing Google's custom AI silicon for next-generation machine learning models. As a Senior Performance Co-Design Engineer, you will focus and conduct LLM Serving Studies. In this role, you will work on analyzing and optimizing the serving performance of emerging models and use cases on our custom hardware. You will also work closely with hardware architects to influence the evolution of Google’s custom ML accelerators. The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide. We're the driving force b [... source excerpt omitted ...] ding job-related skills, experience, and relevant education or training. US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits Learn more about benefits at Google. Responsibilities Conduct comprehensive serving performance studies on (current and emerging) LLMs (1P/3P models). Develop and maintain advanced simulation, profiling, and modeling tools to identify bottlenecks, understand key characteristics and project serving workload performance. Partner with model researchers, software and hardware teams to co-design architectural improvements tailored to Large Language Model (LLM) inference latency and throughput. Drive data-backed decisions that influence the roadm [... source excerpt omitted ...] erience. 5 years of experience in performance modeling/engineering, computer architecture, co-design, or systems engineering. Experience programming in C++ or Python. Preferred qualifications: Master's degree or PhD in Electrical Engineering, Computer Engineering or Computer Science, with an emphasis on computer architecture. Experience with hardware/software co-design problems, especially performance analysis and identification at the pre-silicon stage. Experience enabling and optimizing large-scale ML models (e.g., LLMs, large embedding models). Experience with ML infrastructure, profiling tools, or deep learning inference/serving optimizations. Familiarity with accelerator a
Employer postings · Data from · Sources