About the role
What will you do at Scale AI?
As a Software Engineer on the Machine Learning Infrastructure team, you will
build the "Operating System" for our large-scale GPU clusters. You will
architect a high-performance training platform that handles the immense
complexity of multi-thousand GPU workloads, ensuring every cycle is used
efficiently. Your work directly determines the velocity at which our researchers
can train and iterate on the world’s most advanced models.
The ideal candidate is a systems expert who thrives on solving the
orchestration, networking, and reliability challenges that emerge at massive
scale. You will partner closely with researchers to build a seamless, resilient
environment that transforms raw compute into breakthrough AI.
YOU WILL:
* Architect and scale a multi-tenant orchestration layer that abstracts away
the complexity of GPU clusters, ensuring high utilization and seamless job
recovery.
* Design and implement scheduling primitives to optimize the lifecycle of
training jobs.
* Develop deep observability and automated health-checking into the training
stack to proactively identify and isolate hardware failures
* Evaluate and integrate emerging technologies in the CNCF and AI ecosystem
(e.g. Ray, Kueue), making data-driven build vs. buy decisions that balance
velocity with long-term maintainability.
* Work closely with Finance and Procurement teams to drive our capacity
planning process.
* Participate in our team’s on call process to ensure the availability of our
services.
* Own projects end-to-end, from requirements, scoping, design, to
implementation, in a highly collaborative and cross-functional environment.
IDEALLY YOU'D HAVE:
* 5+ years of experience in backend or infrastructure engineering, with at
least 2 years focused on orchestrating ML workloads at scale (100+ GPU
nodes).
* Strong programming skills in one or more languages (e.g. Python, Go, Rust,
C++)
* Experience with complex compute management systems that cover queueing,
quotas, preemption, and gang scheduling.
* Experience with distributed training infrastructure, such as EFA, Infiniband,
and topology-aware scheduling.
* Experience with distributed storage systems (e.g. Lustre, S3) as they relate
to training throughput
* Expert-level knowledge of Kubernetes internals (Custom Resources, Operators,
Admission Controllers) and how they interact with device plugins for
specialized hardware.
* Familiarity with cloud infrastructure (AWS, GCP) and infrastructure as code
(e.g., Terraform).
* Proven ability to solve complex problems and work independently in
fast-moving environments.
NICE TO HAVES:
* Experience with distributed training techniques such as DeepSpeed, FSDP, etc.
* Experience with the NVIDIA software and hardware stack (CUDA, NCCL)
* Experience with PyTorch
* Familiarity with post-training algorithms such as GRPO, and with
Reinforcement Learning
Compensation
packages at Scale for eligible roles include base salary, equity,
and benefits. The range displayed on each job posting reflects the minimum and
maximum target for new hire salaries for the position and may be inclusive of
several career levels at Scale; it will be determined during the interview
process based on work location and additional factors, including job-related
skills, experience, qualifications, interview performance, and relevant
education or training. Scale employees in eligible roles are also granted equity
based compensation, subject to Board of Director approval. Your recruiter can
share more about the specific salary range for your preferred location during
the hiring process, and confirm whether the hired role will be eligible for
equity grant. You'll also receive benefits including, but not limited to:
comprehensive health, dental and vision coverage, retirement benefits, a
learning and development stipend, and generous PTO. Additionally, this role may
be eligible for additional benefits such as a commuter stipend.
Please reference the job posting's subtitle for where this position will be
located. For pay transparency purposes, the base salary range for this full-time
position in the locations of San Francisco, New York, Seattle is:
$216,000—$270,000 USD
PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering
candidates for the same role. This allows us to ensure a fair and thorough
evaluation of all applicants.
About Us:
At Scale, our mission is to develop reliable AI systems for the world's most
important decisions. Our products provide the high-quality data and full-stack
technologies that power the world's leading models, and help enterprises and
governments build, deploy, and oversee AI applications that deliver real impact.
We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic,
Time Inc., the Government of Qatar, and U.S. government agencies including the
Army and Air Force. We are expanding our team to accelerate the development of
AI applications.
We believe that everyone should be able to bring their whole selves to work,
which is why we are proud to be an inclusive and equal opportunity workplace. We
are committed to equal employment opportunity regardless of race, color,
ancestry, religion, sex, national origin, sexual orientation, age, citizenship,
marital status, disability status, gender identity or Veteran status.
We are committed to working with and providing reasonable accommodations to
applicants with physical and mental disabilities. If you need assistance and/or
a reasonable accommodation in the application or recruiting process due to a
disability, please contact us at [email protected]. Please see the United
States Department of Labor's Know Your Rights poster
[https://www.eeoc.gov/sites/default/files/2023-06/22-088_EEOC_KnowYourRights6.12ScreenRdr.pdf]
for additional information.
We comply with the United States Department of Labor's Pay Transparency
provision.
PLEASE NOTE: We collect, retain and use personal data for our professional
business purposes, including notifying you of job opportunities that may be of
interest and sharing with our affiliates. We limit the personal data we collect
to that which we believe is appropriate and necessary to manage applicants’
needs, provide our services, and comply with applicable laws. Any information we
collect in connection with your application will be treated in accordance with
our internal policies and programs designed to protect personal data. Please see
our privacy policy [https://scale.com/legal/privacy] for additional information.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Neural Data Infrastructure Engineer
Blackrock Neurotech · Salt Lake City, Utah, United States
Role information can change. Confirm current details on the original application page.
