About the role
What will you do at Google?
MINIMUM QUALIFICATIONS:
* Bachelor's degree or equivalent practical experience.
* 8 years of software development experience in C, C++, Go, or Python.
* 5 years of experience testing, and launching software products.
* 5 years of experience building and developing large-scale infrastructure,
distributed systems or networks, or experience with compute technologies,
storage, or hardware architecture.
* 3 years of experience designing, building, and operating large-scale
distributed systems, high-performance networking stacks, or operating system
internals.
PREFERRED QUALIFICATIONS:
* Master’s degree or PhD in Engineering, Computer Science, or a related
technical field.
* 8 years of experience with data structures/algorithms.
* 3 years of experience in a technical leadership role leading project teams
and setting technical direction.
* 3 years of experience working in a complex, matrixed organization involving
cross-functional, or cross-business projects.
* Experience with lower-half server architectures, hardware-adjacent
orchestration, and low-level security implementations.
* Experience with Kubernetes and Google-internal cluster systems, alongside a
proven ability to build telemetry pipelines and monitoring systems for
distributed hardware.
ABOUT THE JOB:
Google's software engineers develop the next-generation technologies that change
how billions of users connect, explore, and interact with information and one
another. Our products need to handle information at massive scale, and extend
well beyond web search. We're looking for engineers who bring fresh ideas from
all areas, including information retrieval, distributed computing, large-scale
system design, networking and data storage, security, artificial intelligence,
natural language processing, UI design and mobile; the list goes on and is
growing every day. As a software engineer, you will work on a specific project
critical to Google’s needs with opportunities to switch teams and projects as
you and our fast-paced business grow and evolve. We need our engineers to be
versatile, display leadership qualities and be enthusiastic to take on new
problems across the full-stack as we continue to push technology forward.
The Emergent AI infrastructure team in Google is looking to build the next
generation of on-prem AI infrastructure to bring the best of Google to empower
Frontier model and AI solution builders to advance AI around the world.
The AI and Infrastructure team is redefining what’s possible. We empower Google
customers with breakthrough capabilities and insights by delivering AI and
Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our
customers include Googlers, Google Cloud customers, and billions of Google users
worldwide.
We're the driving force behind Google's groundbreaking innovations, empowering
the development of our cutting-edge AI models, delivering unparalleled computing
power to global services, and providing the essential platforms that enable
developers to build the future. From software to hardware our teams are shaping
the future of world-leading hyperscale computing, with key teams working on the
development of our TPUs, Vertex AI for Google Cloud, Google Global Networking,
Data Center operations, systems research, and much more.
Individual pay is determined by factors including job-related skills,
experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google
[https://www.google.com/about/careers/applications/benefits/].
RESPONSIBILITIES:
* Lead the design and end-to-end software delivery of specialized AI compute
platforms, ensuring high availability for massive-scale workloads.
* Establish observability and hardening strategies for the lower-half software
stack, including firmware, and hardware qualification.
* Architect robust integration interfaces between custom compute topologies and
industry-standard workload schedulers like Kubernetes and GKE.
* Architect secure boot and cryptographic remote attestation flows for
distributed hardware platforms.
* Partner closely with hardware engineering and chip design teams to influence
future architectures for large-scale deployment.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Data and ML Infrastructure Engineer
HavocAI · United States
Role information can change. Confirm current details on the original application page.
