About the role
What will you do at NVIDIA?
NVIDIA is seeking a principal-level software engineer to build the next generation of our Kubernetes platform. Our teams build foundational capabilities for self-service GPU infrastructure, managed Kubernetes control planes, management of cluster operations, and automation for large-scale AI and improved computational environments.
In this role, you will
- work at the intersection of platform architecture and production-grade Kubernetes lifecycle systems. You will lead technical strategy and execution for systems that make clusters easier to provision, upgrade, operate, and scale across cloud and on-premises environments.
- What you will be doing:
- Lead the architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.
- Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
- Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack.
- Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
- Lead the diagnosis and resolution of complex platform issues spanning infrastructure, runtime, networking, hardware, and operations, improving the scalability, resilience, and operability of systems supporting large-scale AI deployments.
- Influence engineering standards, architectural decisions, and long-term platform strategy.
- Mentor senior engineers and raise the bar for design quality, execution, and engineering rigor across the organization.
