Prepin
Log in
NVIDIA

engineering opportunity

Senior Software Engineer, Resilience Engineering - DGX Cloud

Develop and lead the organization-wide reliability strategy and SLO program for DGX Cloud in a 24/7 environment. Lead high-severity incident response and implement chaos engineering and resilience testing to improve system stability.

Santa Clara, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at NVIDIA?

NVIDIA has been transforming computer graphics, PC gaming, and accelerated

computing for more than 25 years. It’s a unique legacy of innovation that’s

fueled by great technology—and amazing people. Today, we’re tapping into the

unlimited potential of AI to define the next era of computing. An era in which

our GPU acts as the brains of computers, robots, and self-driving cars that can

understand the world. Doing what’s never been done before takes vision,

innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a

diverse, supportive environment where everyone is inspired to do their best

work. Come join the team and see how you can make a lasting impact on the world.

Are you passionate about building world-class reliability systems? Join NVIDIA

as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a

pivotal part of a team that redefines operational excellence. Our team is at the

forefront of redefining how DGX Cloud approaches reliability, making it an

outstanding opportunity to develop strategies and drive innovation. We're

looking for a seasoned engineer with experience in running large-scale systems

and a deep understanding of operational practices. What you'll be doing: Build

org-wide reliability strategy, guiding how NVIDIA matures its operational

practices in a 24/7 environment. Stand up a rigorous SLO program, defining and

maintaining high standards across teams. Lead incident response for high

severity incidents, ensuring low drama and high signal resolution. Build and

improve production code daily, enhancing our data platform and related tooling.

Implement chaos engineering, failure injection, and resilience testing to

elevate our team's standard practices. Improve standards by setting an example

with your hands-on experience and leadership. What we need to see: Deep,

hands-on experience running large-scale production systems with a proven track

record. A detailed understanding of failure modes in large systems, including

cascading dependencies and retry storms. Strong software engineering skills with

current, hands-on experience in Go, Python, or similar languages. Proven

experience in establishing and maintaining an SLO program with operational

rigor. Practical experience in reliability fields such as chaos engineering and

failure injection. The ability to influence across team boundaries through

credibility and expertise. 10+ years of industry experience with a Bachelor's or

Master's degree, or equivalent experience operating systems at scale. Ways to

stand out from the crowd: Experience within a world-class reliability function

like Google SRE or Meta production engineering. Expertise in operating GPU, HPC,

or AI training infrastructure with outstanding failure modes. A track record of

measurable reliability improvements within an organization. Proficiency with

modern observability and operational tools like Prometheus, OpenTelemetry,

Grafana, PagerDuty, and Rootly. Widely considered to be one of the technology

world’s most desirable employers, NVIDIA offers highly competitive salaries and

a comprehensive benefits package. As you plan your future, see what we can offer

to you and your family www.nvidiabenefits.com/ Your base salary will be

determined based on your location, experience, and the pay of employees in

similar positions. The base salary range is 184,000 USD - 287,500 USD for Level

4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for

equity and benefits. Applications for this job will be accepted at least until

June 27, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in

its recruiting processes. NVIDIA is committed to fostering an inclusive work

environment and proud to be an equal opportunity employer. As we highly value

diversity in our current and future employees, we do not discriminate (including

in our hiring and promotion practices) on the basis of race, religion, color,

national origin, gender, gender expression, sexual orientation, age, marital

status, veteran status, disability status or any other characteristic protected

by law. NVIDIA pioneered accelerated computing. Today, our AI infrastructure

powers global intelligence, transforming every industry. Learn more about

NVIDIA.

Which skills does this role require?

Reliability EngineeringSLO Program ManagementIncident ResponseChaos EngineeringFailure InjectionGoPythonProduction EngineeringObservabilityLarge-scale SystemsResilience TestingSoftware EngineeringHPC InfrastructureAI Training InfrastructureData PlatformsOperational ExcellenceDGX CloudSREGPUHPCAI InfrastructurePrometheusOpenTelemetryGrafanaPagerDutyRootlyCascading DependenciesRetry StormsDistributed SystemsReliability StrategyFailure ModesCloud ComputingNVIDIA

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.