About the role
What will you do at NVIDIA?
NVIDIA is looking for a Senior Network Reliability Engineer to support and
maintain our cloud and datacenter network infrastructures. This network serves
the needs across the whole software stack for NVIDIA, from Graphics Drivers to
Autonomous Vehicles and Artificial Intelligence. In this role, the Senior
Network Operations Engineer will remediate critical alerts within defined SLAs,
triage production impacting network incidents, and interact with internal
customers on network related issues. They will also be responsible for engaging
with external vendors to remediate hardware and software issues, and participate
in project related work such as network device upgrades and capacity
augmentations. An ideal candidate will possess a wide range of skills, including
alert monitoring & resolution in large-scale networks and CSP environments,
outstanding troubleshooting skills, understanding of L3 underlay networks, and
network protocol knowledge in large multi-vendor infrastructures. What you will
be doing: Engage in 24/7 global shift rotations to provide remote support for
network repairs and changes while collaborating across teams and updating
customers on status and ticket information. Drive operational improvements in
change management and daily operations by following procedures. Manage and
operate large scale IP network technologies and infrastructures. Utilize your
skills in Peering and Datacenter interconnect technologies: PNI, Transit,
Exchange, Passive DWDM, Wave circuits. Monitor and support the network health of
on-premises and cloud infrastructures. Collaborate and develop workflow
enhancements while documenting best practices. What we need to see: Deep
knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS,
GRE, IPsec, DNS, and MACsec. 5+ years of experience in network operations.
Skilled in network troubleshooting techniques and demonstrating creative
problem-solving abilities. Strong track record of alert response within defined
SLAs and Incident management. Experience with one or more of the following CSP
environments: AWS, Azure, GCP, OCI. Familiarity with Arista, Fortinet and
Juniper. Hands-on experience with contributing to tooling and automation for
provisioning, monitoring, and managing complex network infrastructures.
Bachelor’s degree in Computer Science, related technical field, or equivalent
experience. Excellent verbal and written communication skills. Ways To Stand Out
From The Crowd: Solid understanding of Mellanox/Cumulus OS and Infiniband
technology. Skilled in Unix/Linux system administration, with the ability to
write and understand Python/Shell scripts to improve efficiency in hyperscale
environments. Familiarity with leveraging tools such as Netbox/Nautobot,
Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate
about innovating and investing in ground breaking technologies. NVIDIA is widely
considered to be one of the technology world’s most desirable employers. We have
some of the most forward-thinking and hard-working people in the world working
for us. Are you creative and autonomous? Do you love a challenge?
If so, we want
to hear from you. NVIDIA’s deep learning platforms have made major impact to
various fields is broadly used across leading academic institutions, start-ups,
and industry, including the world’s largest Internet companies. We need
passionate, hard-working and creative people to help us take on more of these
outstanding opportunities in deep learning cloud solutions. Your base salary
will be determined based on your location, experience, and the pay of employees
in similar positions. The base salary range is 136,000 USD - 224,250 USD for
Level 3, and 168,000 USD - 264,500 USD for Level 4. You will also be eligible
for equity and benefits. Applications for this job will be accepted at least
until July 8, 2026. This posting is for an existing vacancy. NVIDIA uses AI
tools in its recruiting processes. NVIDIA is committed to fostering an inclusive
work environment and proud to be an equal opportunity employer. As we highly
value diversity in our current and future employees, we do not discriminate
(including in our hiring and promotion practices) on the basis of race,
religion, color, national origin, gender, gender expression, sexual orientation,
age, marital status, veteran status, disability status or any other
characteristic protected by law. NVIDIA pioneered accelerated computing. Today,
our AI infrastructure powers global intelligence, transforming every industry.
Learn more about NVIDIA.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
AI Outcome Customer Engineer, Forward Deployed Engineering
Google · Atlanta, Georgia, United States
Senior Principal Engineer, Enterprise Networking Architecture, Automation and AI
Equinix · Toronto, Ontario, Canada
AI Evaluations Engineer, US Decision Intelligence
Apple · Cupertino, California, United States
Senior Pre-Sales Solutions Engineer - SIEM/Security Analytics / CTI
Anomali · Boston, Massachusetts, United States
Research Engineer, Responsible Frontier AI Research, DeepMind
Google · New York, New York, United States
VP – Distinguished Engineer of Generative AI Engineering
Slate Auto · United States
Role information can change. Confirm current details on the original application page.
