About the role
What will you do at NVIDIA?
NVIDIA is leading the way in groundbreaking developments in Artificial
Intelligence, High Performance Computing and Visualization. The GPU, our
invention, serves as the visual cortex of modern computers and is at the heart
of our products and services. Our work opens up new universes to explore,
enables amazing creativity and discovery, and powers what were once science
fiction inventions from artificial intelligence to autonomous cars. We are
looking for a motivated Deep Learning engineer to bring advanced communication
technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc.
You will be working with the team that created communication libraries like
NCCL, NVSHMEM & technology like GPUDirect -- for scaling Deep Learning and HPC
applications. Your customers will have diverse multi-GPU demands, ranging from
training on scales up to 100K GPUs to inference down at microsecond latency.
Communication performance between the GPUs has a direct impact on AI
applications. Your work in AI toolkits will make all of those easier for the
community. This is an outstanding opportunity for someone with an AI background
to advance the state of the art in this space. Are you ready to contribute to
the development of innovative technologies and help realize NVIDIA's vision?
What you will be doing: Integrate new communication libraries features in AI
frameworks: from PoC to performance analysis to production Perform deep analysis
of AI workloads and frameworks to identify multi-GPU communication requirements
and opportunities. Collaborate hands-on with teams working on the latest AI
models. Improve AI compilers to hide communications or perform automatic fusion.
Conduct in-depth AI workload performance characterization on multi-GPU clusters.
Design fault-tolerant and elastic solutions for large-scale or dynamic AI
workloads. Author custom communication or fused compute-communication kernels to
showcase ultimate performance on NV platforms. Influence the roadmap of
communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team
across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer
Science, or related field (or equivalent experience) with 5+ software
engineering and HPC/AI experience Development or integration experience with
Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as
TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA
or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or
compiler technologies (e.g. torch.compile) Experience conducting performance
benchmarking on AI clusters. Familiarity with at least one performance profiler
toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI
communication concepts (1-sided v 2-sided communication, elasticity, resiliency,
topology discovery, etc) Adaptability and passion to learn new areas and tools
Flexibility to work and communicate effectively across different teams and
timezones Ways to stand out from the crowd: Experience with parallel programming
on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding
of computer system architecture, HW-SW interactions and operating systems
principles (aka systems software fundamentals) Expertise in one or more of these
areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel
authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute
& communication overlap in distributed runtimes Experience with AI compiler
pattern matching and lowering. Solid understanding of memory hierarchy,
consistency model, and tensor layout Your base salary will be determined based
on your location, experience, and the pay of employees in similar positions. The
base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD -
287,500 USD for Level 4. You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until July 11, 2026. This
posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting
processes. NVIDIA is committed to fostering an inclusive work environment and
proud to be an equal opportunity employer. As we highly value diversity in our
current and future employees, we do not discriminate (including in our hiring
and promotion practices) on the basis of race, religion, color, national origin,
gender, gender expression, sexual orientation, age, marital status, veteran
status, disability status or any other characteristic protected by law. NVIDIA
pioneered accelerated computing. Today, our AI infrastructure powers global
intelligence, transforming every industry. Learn more about NVIDIA.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Staff/Principal AI Transformation Engineer
DiDi Autonomous Driving · San Jose, California, United States
AI Evaluations Engineer, US Decision Intelligence
Apple · Cupertino, California, United States
Research Engineer, Responsible Frontier AI Research, DeepMind
Google · New York, New York, United States
AI Outcome Customer Engineer, Forward Deployed Engineering
Google · Atlanta, Georgia, United States
AI Risk Engineer
Bright Vision Technologies · Columbus, Ohio, United States
Principal Engineer, Content and Generative AI Exploration, Search Platforms
Google · Mountain View, California, United States
Role information can change. Confirm current details on the original application page.