About the role
What will you do at Apple?
The Production Engineering team within the AI and Data Platform (AiDP)
organization manages a wide array of real-time, near real-time, and batch
analytical solutions. These platforms are integral to core business functions
across Apple. These include sales, operations, finance, AppleCare, marketing,
and services, and are instrumental in driving critical, data-driven decisions.
To build these solutions, we leverage a combination of proprietary and leading
open-source technologies such as Kafka, Spark, Iceberg, and Airflow. A key part
of our mission is to enable AI-centric automations that enhance the overall
efficiency and intelligence of the platform. We are looking for passionate
engineers who thrive on solving complex infrastructure challenges at scale, both
on-premises and in the cloud. If you are dedicated to optimizing scalable,
maintainable, and user-friendly systems, you will find compelling opportunities
to make a significant impact at AiDP.
DESCRIPTION
The Service Reliability Engineer (SRE) role within AiDP Production Engineering
is a dynamic position that blends strategic architectural design with hands-on
technical execution. As an SRE, you will be responsible for configuring, tuning,
and ensuring the resilience of complex, multi-tiered systems to achieve optimal
application performance, stability, and availability. Our team manages critical
data pipelines and applications across both bare-metal and cloud computing
platforms, delivering essential data processing for all of Apple’s key business
functions. We operate at an immense scale, handling exabytes of data, petabytes
of memory, and tens of thousands of jobs to enable predictable and performance
data analytics that power features and inform decisions across the company. If
you are passionate about designing, building, and running data infrastructure
that has a direct and significant impact on Apple’s global business operations,
this is the ideal opportunity for you.
MINIMUM QUALIFICATIONS
4+ years experience in cloud-native services, including ETL frameworks like
Apache Spark, and Flink. 4+ years experience in messaging systems (Kafka) and
cloud infrastructure & services, AWS, GCP, Kubernetes. 4+ years of experience in
modern & distributed databases such as Snowflake, Cassandra, SingleStore, and
SAP HANA. 4+ years of programming experience in Python or Java. BS/MS in
computer science or equivalent experience.
PREFERRED QUALIFICATIONS
Solid understanding of system design, data structures, and incident management
best practices. Should be able to understand complex architectures and be
comfortable working with multiple teams. Observability tools (e.g: Prometheus,
Grafana, CloudWatch). Ability to conduct performance analysis and troubleshoot
large scale distributed systems. Should be highly proactive with a keen focus on
improving uptime/availability of our mission critical services. Strong expertise
in troubleshooting complex production issues. Excellent problem solving,
critical thinking, and communication skills. Proven ability to resolve
incidents, perform root cause analysis, and drive system reliability
improvements. Experience using GenAI or automation tools for issue detection,
alerting, or remediation. Experience in data visualization tools such as
Tableau, Business Objects, ThoughtSpot.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Neural Data Infrastructure Engineer
Blackrock Neurotech · Salt Lake City, Utah, United States
Senior Infrastructure Security Engineer
A-Gas · Rhome, Texas, United States
Role information can change. Confirm current details on the original application page.
