Prepin
Log in
Google

engineering opportunity

Senior Software Engineer, AI/ML System Infrastructure

The Senior Software Engineer will drive software development for TPU system control planes and design health management systems based on hardware telemetry. They will also build analytics for failure detection and integrate repair workflows with TPU clusters and cloud infrastructure.

Kirkland, Washington, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Google?

MINIMUM QUALIFICATIONS:

* Bachelor’s degree or equivalent practical experience.

* 5 years of experience working with Go.

* 3 years of experience with developing large-scale infrastructure, distributed

systems or networks, or experience with compute technologies, storage or

hardware architecture.

* 3 years of experience in distributed computing.

* 3 years of experience in infrastructure design.

* 3 years of experience in system architecture.

PREFERRED QUALIFICATIONS:

* Master's degree or PhD in Computer Science or related technical field.

* 5 years of experience with data structures and algorithms.

* 5 years of experience with C and C++.

* 1 year of experience in a technical leadership role.

* Experience developing accessible technologies.

ABOUT THE JOB:

Google's software engineers develop the next-generation technologies that change

how billions of users connect, explore, and interact with information and one

another. Our products need to handle information at massive scale, and extend

well beyond web search. We're looking for engineers who bring fresh ideas from

all areas, including information retrieval, distributed computing, large-scale

system design, networking and data storage, security, artificial intelligence,

natural language processing, UI design and mobile; the list goes on and is

growing every day. As a software engineer, you will work on a specific project

critical to Google’s needs with opportunities to switch teams and projects as

you and our fast-paced business grow and evolve. We need our engineers to be

versatile, display leadership qualities and be enthusiastic to take on new

problems across the full-stack as we continue to push technology forward.

As the Senior Software Engineer, you will drive software development for Tensor

Processing Unit (TPU) system control planes. You will design and implement

health management systems that rely on hardware telemetry. You will build

analytics for detecting hardware problems and develop different health rules

specializing in detecting ICI, OCS, EMD, Tross, and out of band entities failure

modes. You will build algorithms to generate correlated failures and suggest

actions for repair workflows, integrating with both TPU cluster and cloud

infrastructure. You will also build the Diagnoser for the specialized health

rules to maintain and operate these TPU clusters.

Google Cloud accelerates every organization’s ability to digitally transform its

business and industry. We deliver enterprise-grade solutions that leverage

Google’s cutting-edge technology, and tools that help developers build more

sustainably. Customers in more than 200 countries and territories turn to Google

Cloud as their trusted partner to enable growth and solve their most critical

business problems.

Individual pay is determined by factors including job-related skills,

experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google

[https://www.google.com/about/careers/applications/benefits/].

RESPONSIBILITIES:

* Build and develop systems to integrate in continuous running of TPU AI

infrastructure.

* Work cross-functionally to define the requirements for Diagnoser, and how

will the customer benefit from it and define the Critical User Journeys

(CUJs).

* Influence and align the cross-functional teams on roadmap of PodCare and

Diagnoser.

* Deliver high quality code and timely project while working with

cross-functional teams.

* Contribute to CI/CD pipeline and integration and regression test for health

monitoring system.

Which skills does this role require?

GoDistributed SystemsInfrastructure DesignSystem ArchitectureC++Hardware TelemetryCI/CDCloud InfrastructureMachine LearningNetwork EngineeringInfrastructureTensor Processing UnitPodCareDiagnoserHealth MonitoringRegression TestingSoftware EngineeringGCPProduct Strategy

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.