About the role
What will you do at Apple?
At Apple, we focus deeply on our customers’ experience. Apple Ads brings this
same approach to advertising, helping people find exactly what they’re looking
for and helping advertisers grow their businesses. Our technology powers ads and
sponsorships across Apple Services, including the App Store, Apple News, and MLS
Season Pass. Everything we do is designed for trust, connection, and impact: We
respect user privacy, integrate advertising thoughtfully into the experience,
and deliver value for advertisers of all sizes—from small app developers to big,
global brands. Because when advertising is done right, it benefits everyone. The
Site Reliability Engineering team within Apple Ads ensures the reliability,
performance, and availability of ML Platform and Services at scale. The team
partners closely with Ads engineering, data science and ML platform teams to
enable product delivery through design, configuration, and automation of machine
learning infrastructure powering Apple Ads applications. We are looking for a ML
Platform Infrastructure Engineer to help build and evolve the next generation of
Apple Ads machine learning platform — enabling fast, reliable, and scalable
operations across AWS-based environments supporting transactional and analytical
workloads.
DESCRIPTION
As a site reliability engineer in Apple Ads focused on machine learning, you
will own the health, performance, and scalability of large scale infrastructure
powering ML training, inference, serving workloads and associated platform
tooling. Your focus will be on building automation that eliminates manual
processes, improves platform resilience, and enables teams to move faster with
confidence. This is not a DevOps-only or CI/CD-focused role. We are looking for
engineers who build platform solutions, not just configure pipelines.
MINIMUM QUALIFICATIONS
3+ years of experience in internet-facing backend production systems, SRE or ML
Operations focused roles on large scale distributed cloud infrastructure Proven
expertise with AWS-managed infrastructure Familiarity with ML lifecycle and
associated technologies such as NVIDIA Triton, AnyScale Ray, Apache Airflow etc.
Strong programming skills in at least one of: Python, Java, Rust, Go or similar
languages Hands-on experience with Linux systems and deep knowledge of its
internals. Demonstrated experience with Infrastructure as Code, especially
Terraform. Strong foundation in SRE concepts: Monitoring, alerting,
observability, Incident response and root cause analysis, Error budgets,
SLAs/SLOs, and system reliability
PREFERRED QUALIFICATIONS
Built tools or services that automate platform operations, reduce toil, or
improve cost efficiency. Experience managing Kubernetes clusters at scale in
production environments. Hands-on experience troubleshooting distributed systems
under real-world load. Clear communication skills and comfort collaborating
across engineering, infrastructure, and product teams. AWS certifications or
broad experience across multiple AWS services is a plus. Understanding of modern
GPU hardware architectures (such as NVIDIA H100, B200, or GB200, AWS Inferentia
), associated drivers Understanding of high-performance fabrics and network
architecture, power, and thermal limits
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Data and ML Infrastructure Engineer
HavocAI · United States
Data Center Engineer – Windows Infrastructure
Konnect IT Group, Inc. · Chicago, Illinois, United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Role information can change. Confirm current details on the original application page.
