Prepin
Log in
Apple

engineering opportunity

Site Reliability Engineer - ML, Apple Ads

The Site Reliability Engineer will own the health, performance, and scalability of large-scale infrastructure powering ML training, inference, and serving workloads. They will build automation to eliminate manual processes, improve platform resilience, and enable teams to deploy services with confidence.

New York, New York, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

At Apple, we focus deeply on our customers’ experience. Apple Ads brings this

same approach to advertising, helping people find exactly what they’re looking

for and helping advertisers grow their businesses. Our technology powers ads and

sponsorships across Apple Services, including the App Store, Apple News, and MLS

Season Pass. Everything we do is designed for trust, connection, and impact: We

respect user privacy, integrate advertising thoughtfully into the experience,

and deliver value for advertisers of all sizes—from small app developers to big,

global brands. Because when advertising is done right, it benefits everyone. The

Site Reliability Engineering team within Apple Ads ensures the reliability,

performance, and availability of ML Platform and Services at scale. The team

partners closely with Ads engineering, data science and ML platform teams to

enable product delivery through design, configuration, and automation of machine

learning infrastructure powering Apple Ads applications. We are looking for a ML

Platform Infrastructure Engineer to help build and evolve the next generation of

Apple Ads machine learning platform — enabling fast, reliable, and scalable

operations across AWS-based environments supporting transactional and analytical

workloads.

DESCRIPTION

As a site reliability engineer in Apple Ads focused on machine learning, you

will own the health, performance, and scalability of large scale infrastructure

powering ML training, inference, serving workloads and associated platform

tooling. Your focus will be on building automation that eliminates manual

processes, improves platform resilience, and enables teams to move faster with

confidence. This is not a DevOps-only or CI/CD-focused role. We are looking for

engineers who build platform solutions, not just configure pipelines.

MINIMUM QUALIFICATIONS

3+ years of experience in internet-facing backend production systems, SRE or ML

Operations focused roles on large scale distributed cloud infrastructure Proven

expertise with AWS-managed infrastructure Familiarity with ML lifecycle and

associated technologies such as NVIDIA Triton, AnyScale Ray, Apache Airflow etc.

Strong programming skills in at least one of: Python, Java, Rust, Go or similar

languages Hands-on experience with Linux systems and deep knowledge of its

internals. Demonstrated experience with Infrastructure as Code, especially

Terraform. Strong foundation in SRE concepts: Monitoring, alerting,

observability, Incident response and root cause analysis, Error budgets,

SLAs/SLOs, and system reliability

PREFERRED QUALIFICATIONS

Built tools or services that automate platform operations, reduce toil, or

improve cost efficiency. Experience managing Kubernetes clusters at scale in

production environments. Hands-on experience troubleshooting distributed systems

under real-world load. Clear communication skills and comfort collaborating

across engineering, infrastructure, and product teams. AWS certifications or

broad experience across multiple AWS services is a plus. Understanding of modern

GPU hardware architectures (such as NVIDIA H100, B200, or GB200, AWS Inferentia

), associated drivers Understanding of high-performance fabrics and network

architecture, power, and thermal limits

Which skills does this role require?

Site Reliability EngineeringMachine Learning InfrastructurePythonJavaRustGoLinuxTerraformNVIDIA TritonAnyScale RayApache AirflowObservabilityInfrastructure as CodeMachine LearningMLOpsCloud InfrastructureMonitoringIncident ResponseSLAsSLOsGPU ArchitectureAutomationScalabilityPerformance EngineeringData ScienceApple AdsAirflow

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.