Prepin
Log in
Apple

engineering opportunity

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

You will operate and support a multi-cloud data platform, managing services like Spark, Flink, and LLM-based agent platforms. Additionally, you will lead incident response and collaborate with developers to ensure the reliability and scalability of infrastructure across AWS, GCP, and on-premise Kubernetes.

Austin, Texas, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

The Apple Services Engineering team (ASE) is one of the most exciting examples

of Apple's long-held passion for combining art and technology. These are the

people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple

Books — at extensive scale, meeting high expectations to deliver a huge variety

of entertainment in over 35 languages to more than 150 countries. Within ASE,

the Apple Data Platform SRE team keeps a massive, multi-cloud platform running

for thousands of internal engineers building the next generation of data and AI

products at Apple. We sit at the intersection of infrastructure, automation, and

customer success — running incident response, providing hands-on support to

internal teams, and partnering with developers to make cutting-edge services

like Spark, Flink, Airflow, Ray, Notebooks, and LLM-based agent platforms

reliable at scale across AWS, GCP, and on-premise Kubernetes.

DESCRIPTION

This is a rare opportunity to build deep expertise across one of the most

technically diverse platforms at Apple — while specialising in the cloud

infrastructure that underpins all of it. As an SRE on Apple Data Platform,

you'll operate and support the team's full portfolio, from big data pipelines to

ML/AI platform services, and grow into the team's go-to expert for multi-cloud

infrastructure — including AWS core services (IAM, EKS, RDS, S3, VPC networking,

autoscaling, EBS), Kubernetes administration at scale, and the

Infrastructure-as-Code and GitOps tooling that keeps it all reconciled and

reliable. You'll be the person other engineers turn to when an IAM policy

misfires, a cluster hits a scheduling wall, or a GitOps reconciliation drifts

out of sync — and the driving force behind making those failure modes rarer over

time. We're looking for a self-motivated engineer who thrives on ownership —

someone who wants a set of services to call their own, the autonomy to drive

their reliability roadmap, and the collaborative instinct to keep that work

aligned with the team's broader direction. If you love going deep on cloud and

Kubernetes internals, enjoy being the trusted expert customers turn to, and want

a front-row seat to Apple Data Platform's multi-cloud evolution, this role

offers real room to grow your scope and impact over time.

MINIMUM QUALIFICATIONS

Minimum Qualifications

  • * Bachelor's Degree in Computer Science, an
  • engineering-related field, or equivalent related experience. 1-4 years in a Site
  • Reliability Engineering, DevOps, or Infrastructure-focused role. Proficient in
  • Python; working knowledge of Golang a plus. Strong hands-on AWS experience: IAM
  • (roles, policies, permission boundaries, KMS), EKS, RDS, S3, VPC
  • networking/endpoints, autoscaling groups, EBS. Kubernetes administration
  • experience — RBAC, node/pod scheduling, autoscalers, PriorityClasses/PDBs, and
  • troubleshooting cluster-wide disruptions. Strong communication skills and
  • composure under pressure during incidents. Solid grounding in SRE principles,
  • with prior on-call or production-support experience.
  • PREFERRED QUALIFICATIONS
  • Experience with Infrastructure-as-Code (Crossplane and/or Terraform), including
  • debugging state drift and composition/controller issues. Experience with GitOps
  • workflows (Flux or similar) — HelmRepository/reconciliation troubleshooting and
  • Helm chart deployment. Multi-cloud exposure (GCP) — parity and migration
  • scenarios are emerging areas of focus. Experience with Splunk for log pipeline
  • debugging (e.g., fluent-bit). Familiarity with Spark/Flink running on Kubernetes
  • (executor scheduling, node affinity). Comfort with GitHub PR review workflows in
  • an infrastructure-as-code / GitOps context. A track record of automating manual
  • operations through scripting or tooling. Intellectual curiosity and a drive to
  • keep learning — for yourself, your team, and the org.

Which skills does this role require?

PythonGolangVPC NetworkingSite Reliability EngineeringDevOpsAirflowRayLLMAutomationIncident ResponseData PlatformMachine LearningAINode.jsGoLLMsProduct Strategy

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.