Skip to content
NEXTMOVEFDE careers · United States

Senior Software Engineer, AI/ML System Infrastructure

AI summary of the role

Senior Software Engineer on Google Cloud's TPU infrastructure team, responsible for building health management systems and diagnostics for TPU clusters.

What you’ll do

  • Build and develop systems to integrate in continuous running of TPU AI infrastructure.
  • Work cross-functionally to define requirements for Diagnoser and define Critical User Journeys (CUJs).
  • Influence and align cross-functional teams on roadmap of PodCare and Diagnoser.
  • Deliver high quality code and timely project while working with cross-functional teams.

What you’ll bring

  • Bachelor's degree or equivalent practical experience.
  • 5 years of experience working with Go.
  • 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
  • 3 years of experience in distributed computing.

Technologies

Go · C · C++ · TPU · Tensor Processing Unit · hardware telemetry · distributed systems · CI/CD · infrastructure design · system architecture

About Google

Builds global consumer, ads, cloud, developer and AI platforms spanning Search, YouTube, Android, Workspace and Gemini.

Public

Source and classification

Internal deployment & tooling · Evidence for this classification:

full-stack as we continue to push technology forward. As the Senior Software Engineer, you will drive software development for Tensor Processing Unit (TPU) system control planes. You will design and implement health management systems that rely on hardware telemetry. You will build analytics for detecting hardware problems and develop different health rules specializing in detecting ICI, OCS, EMD, Tross, and out of band entities failure modes. You will build algorithms to generate correlated failures and suggest actions for repair workflows, integrating with both TPU cluster and cloud infrastructure. You will also build the Diagnoser for the specialized health rules to maintain and operate these TPU clusters. Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge
More from the job description

About the job Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward. As the Senior Software Engineer, you will drive software development for Tensor Processing Unit (TPU) system control planes. You will design and implement health management systems that rely on hardware telemetry. You will build analytics for detecting hardware problems and develop different health rules specializing in detecting ICI, OCS, EMD, Tross, and out of band entities failure modes. You will build algorithms to generate correlated failures [... source excerpt omitted ...] y transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems. Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits Learn more about benefits at Google. Responsibilities Build and develop systems to integrate in continuous running of TPU AI infrastructure. Work cross-f [... source excerpt omitted ...] er high quality code and timely project while working with cross-functional teams. Contribute to CI/CD pipeline and integration and regression test for health monitoring system. Qualifications Minimum qualifications: Bachelor’s degree or equivalent practical experience. 5 years of experience working with Go. 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture. 3 years of experience in distributed computing. 3 years of experience in infrastructure design. 3 years of experience in system architecture. Preferred qualifications: Master's degree or PhD in Compu

How jobs are selected

Employer postings · Data from · Sources