About the role
What will you do at NVIDIA?
We are looking for a Senior Deep Learning Software Engineer to design and build
our automated inference and deployment solution. As part of the team, you will
be instrumental in defining a scalable architecture for DL inference with
emphasis on ease-of-use and compute efficiency. Your work will span multiple
layers of the DL deployment stack, encompassing developing features in
high-level frameworks like PyTorch and JAX, designing and implementing a
high-performance execution environment, low-level GPU optimizations and
developing custom GPU kernels in CUDA and/or Triton. This is an exceptional
opportunity for passionate software engineers straddling the boundaries of
research and engineering, with a strong background in both machine learning
fundamentals and software architecture & engineering. What you’ll be doing: Play
a pivotal role in defining of a modular, scalable platform to seamlessly bridge
training and deployment workflows—enabling tight integration of deployment
tooling with training frameworks such as Megatron and Nemo Leverage and build
upon the torch 2.0 ecosystem (TorchDynamo, torch.export, torch.compile, etc...)
to analyze and extract standardized model graph representation from arbitrary
torch models for our automated deployment solution. Develop support for
inference optimization techniques such as speculative decoding and LoRA.
Collaborate with teams across NVIDIA to use performant kernel implementations
within the automated deployment solution. Analyze and profile GPU kernel-level
performance to identify hardware and software optimization opportunities.
Continuously innovate on the inference performance to ensure NVIDIA's inference
software solutions (TRT, TRT-LLM, TRT Model Optimizer) can maintain and increase
its leadership in the market. What we need to see: Masters, PhD, or equivalent
experience in Computer Science, AI, Applied Math, or related field. 8+ years of
relevant work or research experience in Deep Learning. Excellent software design
skills, including debugging, performance analysis, and test design. Strong
proficiency in Python, PyTorch, and related ML tools. Strong algorithms and
programming fundamentals. Good written and verbal communication skills and the
ability to work independently and collaboratively in a fast-paced environment.
Ways to stand out from the crowd: Contributions to PyTorch, JAX, or other
Machine Learning Frameworks. Knowledge of GPU architecture and compilation
stack, and capability of understanding and debugging end-to-end performance.
Familiarity with NVIDIA's deep learning SDKs such as TensorRT. Prior experience
in writing high-performance GPU kernels for machine learning workloads in
frameworks such as CUDA, CUTLASS, or Triton. Increasingly known as “the AI
computing company” and widely considered to be one of the technology world’s
most desirable employers. Are you creative, motivated, and love a challenge? If
so, we want to hear from you! Come, join our model optimization group, where you
can help build real-time, cost-effective computing platforms driving our success
in this exciting and rapidly-growing field. #LI-Hybrid Your base salary will be
determined based on your location, experience, and the pay of employees in
similar positions. The base salary range is 224,000 USD - 356,500 USD. You will
also be eligible for equity and benefits. Applications for this job will be
accepted at least until April 28, 2026. This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to
fostering a diverse work environment and proud to be an equal opportunity
employer. As we highly value diversity in our current and future employees, we
do not discriminate (including in our hiring and promotion practices) on the
basis of race, religion, color, national origin, gender, gender expression,
sexual orientation, age, marital status, veteran status, disability status or
any other characteristic protected by law. NVIDIA pioneered accelerated
computing. Today, our AI infrastructure powers global intelligence, transforming
every industry. Learn more about NVIDIA.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior SRAM Circuit Design Engineer - AI & HPC (7708)
TSMC · San Jose, California, United States
AI Workflow Engineer
Scout Motors Inc. · Charlotte, North Carolina, United States
AI Evaluations Engineer, US Decision Intelligence
Apple · Cupertino, California, United States
Research Engineer, Responsible Frontier AI Research, DeepMind
Google · New York, New York, United States
AI Outcome Customer Engineer, Forward Deployed Engineering
Google · Atlanta, Georgia, United States
AI Risk Engineer
Bright Vision Technologies · Columbus, Ohio, United States
Role information can change. Confirm current details on the original application page.