Prepin
Log in
Apple

engineering opportunity

Site Reliability Engineer, Apple Data Platform - AI/ML Platform

You will operate and support a multi-cloud platform for AI/ML products, ensuring the reliability of services like Spark, Flink, and Ray. You will also lead incident response and partner with developers to build and maintain scalable infrastructure for data and AI workflows.

Austin, Texas, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

The Apple Services Engineering team (ASE) is one of the most exciting examples

of Apple's long-held passion for combining art and technology. These are the

people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple

Books — at extensive scale, meeting high expectations to deliver a huge variety

of entertainment in over 35 languages to more than 150 countries. Within ASE,

the Apple Data Platform SRE team keeps a massive, multi-cloud platform running

for thousands of internal engineers building the next generation of data and AI

products at Apple. We sit at the intersection of infrastructure, automation, and

customer success — running incident response, providing hands-on support to

internal teams, and partnering with developers to make cutting-edge services

like Spark, Flink, Airflow, Ray, Notebooks and LLM-based agent platforms

reliable at scale.

DESCRIPTION

This is a rare opportunity to build deep expertise across one of the most

technically diverse platforms at Apple — while specialising in an area that's

shaping the future of how Apple builds and operates AI. As an SRE on Apple Data

Platform, you'll operate and support the team's full portfolio, from big data

pipelines to multi-cloud infrastructure, and grow into the team's go-to expert

for ML/AI platform services — including Ray training and serving, LangGraph

agent deployments, RAG architectures, embeddings platforms, and vector store

platforms. You won't be building the models yourself, but you'll be the

infrastructure backbone behind the teams who do — keeping their services,

pipelines, and platforms running flawlessly in production so they can focus on

innovation. We're looking for a self-motivated engineer who thrives on ownership

— someone who wants a set of services to call their own, the autonomy to drive

their reliability roadmap, and the collaborative instinct to keep that work

aligned with the team's broader direction. If you love solving hard operational

problems, enjoy being the trusted expert customers turn to, and want a front-row

seat to Apple's ML/AI infrastructure evolution, this role offers real room to

grow your scope and impact over time.

MINIMUM QUALIFICATIONS

Minimum Qualifications

  • * Bachelor's Degree in Computer Science, an
  • engineering-related field, or equivalent related experience. 1-4 years in a Site
  • Reliability Engineering, DevOps, or Infrastructure-focused role. Proficient in
  • Python; working knowledge of Golang a plus. Experience with Kubernetes and at
  • least one major cloud provider (AWS or GCP). Exposure to operating or supporting
  • ML pipelines, model-serving infrastructure, or LLM-based systems in production.
  • Strong communication skills and composure under pressure during incidents. Solid
  • grounding in SRE principles, with prior on-call or production-support
  • experience.
  • PREFERRED QUALIFICATIONS
  • Hands-on experience operating or supporting Ray (training/serving), LangGraph or
  • similar agent orchestration frameworks, RAG architectures, embeddings platforms,
  • or vector store platforms. Familiarity with MCP-based tooling and ML
  • lifecycle/dataset management systems. Experience with S3 and cloud
  • storage/networking fundamentals. Familiarity with observability tooling:
  • Prometheus, Grafana, Splunk, PagerDuty. Working knowledge of CI/CD pipelines and
  • deployment workflows. Deep understanding of one or more Big Data technologies
  • (Spark, Flink, Airflow, Trino, Notebooks). A track record of automating manual
  • operations through scripting or tooling. Intellectual curiosity and a drive to
  • keep learning — for yourself, your team, and the org.

Which skills does this role require?

PythonSite Reliability EngineeringCloud InfrastructureMachine LearningData PipelinesGolangIncident ResponseAutomationVector StoresData PlatformAI/MLModel ServingCloud ComputingGoLLMsProduct Strategy

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.