About the role
What will you do at Apple?
Scaling machine learning workloads across thousands of accelerators creates
challenges that few engineers ever encounter. In Apple’s Machine Learning
Platform Technologies organization, we build the infrastructure that powers
large-scale ML training and inference workloads, bringing together expertise in
distributed systems, machine learning infrastructure, and high-performance
computing.
DESCRIPTION
As a senior engineer on the ML Compute Capacity team, you will design, build,
and operate the production systems that ensure compute resources are optimally
distributed throughout the company. You'll work across the stack — from data
pipelines and backend services to APIs and interactive frontends — developing
telemetry systems, optimization algorithms, policies, and intuitive tools for
managing demand and improving efficiency across Apple's largest accelerator
fleet. Our small, nimble team works in a high-autonomy, fast-paced environment,
and we're passionate about digging into data patterns, laying out the
performance characteristics of an entire distributed system, and knowledge
sharing. If the opportunity to own and operate services that scale, stay highly
available, and "just work" excites you, then please reach out to us!
MINIMUM QUALIFICATIONS
5+ years of experience in relevant areas Proficiency in Python for production
backend and data engineering work Experience building data pipelines and
crafting robust queries over large-scale, multi-source data (e.g., Trino,
PostgreSQL, Elasticsearch) Experience designing and building RESTful APIs and
working with cloud storage technologies Experience with modern web frameworks
like React Experience with observability tools (e.g., Prometheus, Grafana) or
equivalent monitoring systems Excellent problem-framing and problem-solving
skills Strong CS fundamentals Bachelor's degree or higher in Engineering,
Mathematics, Economics, or a related quantitative field
PREFERRED QUALIFICATIONS
Experience operating Kubernetes at production scale — including scheduling,
resource management, and cluster debugging Familiarity with accelerator
utilization patterns across ML training and inference Strong interest with
capacity planning, cost attribution, or FinOps systems
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Sr. Full Stack Engineer - NextJS, AWS, AI
Perficient · Creve Coeur, Missouri, United States
Senior Full Stack Product Engineer
ActiveProspect · Austin, Texas, United States
Senior Full Stack Software Engineer - Java/AWS
The Depository Trust & Clearing Corporation (DTCC) · Tampa, Florida, United States
Full Stack Engineer
LightFeather · United States
Senior Full Stack Engineer
Fidelity Investments · Greenwood Village, Colorado, United States
Senior Full Stack Engineer - AI Platform
Workday · Atlanta, Georgia, United States
Role information can change. Confirm current details on the original application page.
