Prepin
Log in
NVIDIA

engineering opportunity

Senior Network Reliability Engineer - DGX Cloud

Maintain and support cloud and datacenter network infrastructures, remediating critical alerts and triaging production incidents within SLAs. Collaborate with external vendors and internal teams to perform network upgrades and capacity augmentations in a 24/7 global shift rotation.

Santa Clara, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at NVIDIA?

NVIDIA is looking for a Senior Network Reliability Engineer to support and

maintain our cloud and datacenter network infrastructures. This network serves

the needs across the whole software stack for NVIDIA, from Graphics Drivers to

Autonomous Vehicles and Artificial Intelligence. In this role, the Senior

Network Operations Engineer will remediate critical alerts within defined SLAs,

triage production impacting network incidents, and interact with internal

customers on network related issues. They will also be responsible for engaging

with external vendors to remediate hardware and software issues, and participate

in project related work such as network device upgrades and capacity

augmentations. An ideal candidate will possess a wide range of skills, including

alert monitoring & resolution in large-scale networks and CSP environments,

outstanding troubleshooting skills, understanding of L3 underlay networks, and

network protocol knowledge in large multi-vendor infrastructures. What you will

be doing: Engage in 24/7 global shift rotations to provide remote support for

network repairs and changes while collaborating across teams and updating

customers on status and ticket information. Drive operational improvements in

change management and daily operations by following procedures. Manage and

operate large scale IP network technologies and infrastructures. Utilize your

skills in Peering and Datacenter interconnect technologies: PNI, Transit,

Exchange, Passive DWDM, Wave circuits. Monitor and support the network health of

on-premises and cloud infrastructures. Collaborate and develop workflow

enhancements while documenting best practices. What we need to see: Deep

knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS,

GRE, IPsec, DNS, and MACsec. 5+ years of experience in network operations.

Skilled in network troubleshooting techniques and demonstrating creative

problem-solving abilities. Strong track record of alert response within defined

SLAs and Incident management. Experience with one or more of the following CSP

environments: AWS, Azure, GCP, OCI. Familiarity with Arista, Fortinet and

Juniper. Hands-on experience with contributing to tooling and automation for

provisioning, monitoring, and managing complex network infrastructures.

Bachelor’s degree in Computer Science, related technical field, or equivalent

experience. Excellent verbal and written communication skills. Ways To Stand Out

From The Crowd: Solid understanding of Mellanox/Cumulus OS and Infiniband

technology. Skilled in Unix/Linux system administration, with the ability to

write and understand Python/Shell scripts to improve efficiency in hyperscale

environments. Familiarity with leveraging tools such as Netbox/Nautobot,

Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate

about innovating and investing in ground breaking technologies. NVIDIA is widely

considered to be one of the technology world’s most desirable employers. We have

some of the most forward-thinking and hard-working people in the world working

for us. Are you creative and autonomous? Do you love a challenge?

If so, we want

to hear from you. NVIDIA’s deep learning platforms have made major impact to

various fields is broadly used across leading academic institutions, start-ups,

and industry, including the world’s largest Internet companies. We need

passionate, hard-working and creative people to help us take on more of these

outstanding opportunities in deep learning cloud solutions. Your base salary

will be determined based on your location, experience, and the pay of employees

in similar positions. The base salary range is 136,000 USD - 224,250 USD for

Level 3, and 168,000 USD - 264,500 USD for Level 4. You will also be eligible

for equity and benefits. Applications for this job will be accepted at least

until July 8, 2026. This posting is for an existing vacancy. NVIDIA uses AI

tools in its recruiting processes. NVIDIA is committed to fostering an inclusive

work environment and proud to be an equal opportunity employer. As we highly

value diversity in our current and future employees, we do not discriminate

(including in our hiring and promotion practices) on the basis of race,

religion, color, national origin, gender, gender expression, sexual orientation,

age, marital status, veteran status, disability status or any other

characteristic protected by law. NVIDIA pioneered accelerated computing. Today,

our AI infrastructure powers global intelligence, transforming every industry.

Learn more about NVIDIA.

Which skills does this role require?

TCP/IPBGPOSPFMPLSIS-ISVxLANEVPNNetwork TroubleshootingIncident ManagementCSP EnvironmentsAristaFortinetJuniperPythonShell ScriptingInfinibandDGX CloudNetwork Reliability EngineeringQoSGREIPsecDNSMACsecAWSAzureGCPOCIMellanoxCumulus OSUnixLinuxShellNetboxNautobotPrometheus

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.