Prepin
Log in
Google

engineering opportunity

Software Engineer III, TPU Performance, Hardware and Software Codesign

You will bridge the gap between ML workloads and custom TPU hardware by analyzing and optimizing distributed systems and compiler architectures. Additionally, you will drive full-stack hardware-software co-design to improve performance for business-critical production models.

Sunnyvale, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Google?

MINIMUM QUALIFICATIONS:

* Bachelor’s degree or equivalent practical experience.

* 2 years of experience with software development in one or more programming

languages.

* 2 years of experience with performance, large-scale systems data analysis,

visualization tools, or debugging.

* 2 years of experience with computer architecture, performance analysis, and

performance modeling.

PREFERRED QUALIFICATIONS:

* Master's degree or PhD in Computer Science or related technical fields.

* 2 years of experience with data structures and algorithms.

* Experience developing accessible technologies.

ABOUT THE JOB:

Google's software engineers develop the next-generation technologies that change

how billions of users connect, explore, and interact with information and one

another. Our products need to handle information at massive scale, and extend

well beyond web search. We're looking for engineers who bring fresh ideas from

all areas, including information retrieval, distributed computing, large-scale

system design, networking and data storage, security, artificial intelligence,

natural language processing, UI design and mobile; the list goes on and is

growing every day. As a software engineer, you will work on a specific project

critical to Google’s needs with opportunities to switch teams and projects as

you and our fast-paced business grow and evolve. We need our engineers to be

versatile, display leadership qualities and be enthusiastic to take on new

problems across the full-stack as we continue to push technology forward.

In this role, you will

  • bridge the gap between ML workloads and custom Tensor
  • Processing Unit (TPU) hardware. You will analyze and optimize how distributed
  • systems, compiler architectures—such as Accelerated Linear Algebra (XLA)—and
  • emerging software abstractions, such as Compound AI and multi-step agentic
  • systems, execute across Google’s AI infrastructure. You will collaborate with
  • various product area architects within Google, such as Google Cloud and YouTube,
  • and external customers to systematically onboard novel workloads with engaged
  • performance. Your optimizations will directly drive TPU adoption, secure
  • pre-sales engagements, and shape our future ML infrastructure roadmap.
  • Google Cloud accelerates every organization’s ability to digitally transform its
  • business and industry. We deliver enterprise-grade solutions that leverage
  • Google’s cutting-edge technology, and tools that help developers build more
  • sustainably. Customers in more than 200 countries and territories turn to Google
  • Cloud as their trusted partner to enable growth and solve their most critical
  • business problems.Individual pay is determined by factors including job-related
  • skills, experience, and relevant education or training.
  • US: $147000 - $210000 (USD) + 15% bonus target + equity + benefits
  • Learn more about benefits at Google
  • [https://www.google.com/about/careers/applications/benefits/].
  • RESPONSIBILITIES:
  • * Develop and scale benchmarking and workload characterization strategies to
  • enable fast grounding-to-silicon, root-cause performance analysis, and TPU
  • mapping optimization.
  • * Drive full-stack hardware-software co-design to optimize current and future
  • ML accelerator architectures for business-critical production models (e.g.,
  • LLMs and embedding models).
  • * Partner with Product Areas (e.g., YouTube and Ads) to scale key workload
  • pipelines efficiently (Perf/$/Watts) during TPU Pilot and General
  • Availability (GA) transitions.
  • * Build and upgrade compiler-aware simulator tools, hardware cost-models, and
  • performance-ladder pathways to baseline and project physical silicon
  • capabilities.
  • * Distill complex performance analyses and hardware trade-offs into
  • presentations to guide TPU roadmap decision-making in core leadership forums
  • (e.g., ArchForums, NPI, BCR reviews, and TdJs).

Which skills does this role require?

Software developmentPerformance analysisLarge-scale systemsData analysisDebuggingComputer architecturePerformance modelingMachine learningTensor Processing UnitDistributed systemsCompiler architectureHardware-software co-designPerformance AnalysisHardware-Software CodesignMachine LearningDistributed SystemsCompiler ArchitectureLLMsEmbedding ModelsBenchmarkingWorkload CharacterizationGoogle CloudPerformance ModelingSiliconCompound AIAgentic SystemsGCPProduct Strategy

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.