About the role
What will you do at Nscale?
About Nscale
Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work.
The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take
real surface area: the automation and tooling other engineers depend on, and the reliability of
production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll
be expected to make the systems you touch quieter over time.
What You'll Do
- • Build and own the automation and tooling that keeps the platform running; treat operational toil as
- a bug to be fixed, not a fact of life.
- • Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
- • Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
- incident reviews that actually change the system.
- • Investigate performance and reliability problems across Linux, networking, and distributed services,
- then fix them at the source.
- • Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
- stack.
- • Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring
- • 3-6 years in SRE, systems engineering, or software engineering, including time running production in
- a data center or cloud environment.
- • Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
- away.
- • Solid command of Linux, networking fundamentals, and distributed systems.
- • A track record of troubleshooting live production issues and owning the fix through to the retro.
- • Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
- • Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
- asked.
Nice to Have
- • Experience with AI or GPU workloads, or high-performance computing (HPC).
- • Familiarity with high-performance networking (InfiniBand, RDMA).
- • Kubernetes, plus virtualized or bare-metal environments.
- On-Call and Pace
- A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
- and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth
- acting on rather than just an interruption. The goal is to make the systems quieter over time, so each
- rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in
- better shape than you found it, you'll do well here.
- What We Offer
- • Competitive base plus equity, reviewed every 12 months.
- • Real scope early, and a progression plan built around the skills you want to sharpen.
- • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
- your day.
- Salary Range
- $130,000 - $200,000 USD. Actual compensation varies with skill set, experience, and location, and the
- role may be eligible for bonus and equity.
- Equal Opportunities Statement
- At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
- If there’s anything we can do to accommodate your specific situation, please let us know.
- For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Data and ML Infrastructure Engineer
HavocAI · United States
Infrastructure Engineer I - Data Protection
USAA · Tampa, Florida, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
Data Center Engineer – Windows Infrastructure
Konnect IT Group, Inc. · Chicago, Illinois, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Role information can change. Confirm current details on the original application page.
