About the role
What will you do at Google?
MINIMUM QUALIFICATIONS:
* Bachelor's degree in Computer Science, a related technical field, or
equivalent practical experience.
* 8 years of experience with software development in one or more programming
languages (e.g., Python, C, C++, Java, JavaScript).
PREFERRED QUALIFICATIONS:
* Experience collaborating effectively across organizational boundaries,
building relationships, and importing and exporting ideas to achieve broad
organizational goals.
* Experience leading technical strategy and influencing engineering direction
across multiple teams.
* Experience leading cross-organizational technical programs or initiatives at
scale.
* Experience with agentic systems and using AI to support production
operations, including familiarity with LLM-based tooling, AI-driven
automation, or intelligent incident response systems.
* Ability to understand complex relationship between the organization and its
environment, identify connections, adopt different perspectives and quickly
respond to changing circumstances in a strategic way.
ABOUT THE JOB:
Site Reliability Engineering (SRE) combines software and systems engineering to
build and run large-scale, massively distributed, fault-tolerant systems. SRE
ensures that Google's services—both our internally critical and our
externally-visible systems—have reliability, uptime appropriate to users' needs
and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful
eye on our systems capacity and performance.
Much of our software development focuses on optimizing existing systems,
building infrastructure and eliminating work through automation. On the SRE
team, you’ll have the opportunity to manage the complex challenges of scale
which are unique to Google, while using your expertise in coding, algorithms,
complexity analysis and large-scale system design.
SRE's culture of intellectual curiosity, problem solving and openness is key to
its success. Our organization brings together people with a wide variety of
backgrounds, experiences and perspectives. We encourage them to collaborate,
think big and take risks in a blame-free environment. We promote self-direction
to work on meaningful projects, while we also strive to create an environment
that provides the support and mentorship needed to learn and grow.
To learn more: check out our books on Site Reliability Engineering
[https://landing.google.com/sre/book.html] or read a career profile
[https://careers.google.com/stories/site-reliability-engineering-profile-google/]
about why a Software Engineer chose to join SRE.
In this role, you will
- develop, design, and drive technical engineering projects
- and decisions for Apps, Video and Display (AViD) infrastructure and production
- systems. You will mentor and lead other executive level engineers within AViD
- SRE and across Ads to achieve our aligned mission.
- Individual pay is determined by factors including job-related skills,
- experience, and relevant education or training.
- US: $262000 - $364000 (USD) + 25% bonus target + equity + benefits
- Learn more about benefits at Google
- [https://www.google.com/about/careers/applications/benefits/].
- RESPONSIBILITIES:
- * Ensure that Google Ads services within the AViD ecosystem have reliability
- and uptime appropriate to users' needs with a fast rate of improvement, while
- keeping an ever-watchful eye on capacity and performance.
- * Build creative engineering solutions to operations and infrastructure
- problems, including AI-powered automation and agentic workflows for
- production operations.
- * Lead and contribute to the cross-SRE AI Ops program, driving the strategic
- adoption of AI/ML tools to improve incident response, reduce toil, and
- enhance service reliability across Ads SRE.
- * Focus on optimizing existing systems, build infrastructure, and eliminate
- work through automation and AI-driven operational improvements.
- * Determine how our systems relate to each other and use a breadth of tools and
- approaches — including emerging AI capabilities — to solve a broad spectrum
- of problems.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Data and ML Infrastructure Engineer
HavocAI · United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Role information can change. Confirm current details on the original application page.
