About the role
What will you do at Apple?
At Apple, great ideas turn into phenomenal products, services, and customer
experiences at a pace few companies can match. We are seeking a highly
experienced ML Engineer to build, deploy, optimize and operationalize Small and
Large Language Model (LLM)-based applications, with a strong emphasis on
MLOps/LLMOps and scalable production systems.
DESCRIPTION
As an AI Engineer on our team, you will own the infrastructure and tooling that
let LLM-powered features ship reliably at Apple scale: the CI/CD pipelines and
serving infrastructure that get a model into production, and the observability,
versioning, and governance that keep it trustworthy once it's there. You'll work
across the full model lifecycle, from experimentation and fine-tuning through
deployment, monitoring, and retirement. That ownership extends to the data
feeding these systems and the infrastructure serving them. You'll build
pipelines that ingest and enrich multimodal data through feature stores and
lineage-tracked storage, deploy and operate services on cloud-native
infrastructure such as Kubernetes, and expose them through well-modeled APIs.
You'll also optimize models for production through quantization, distillation,
and compilation, and implement the governance workflows, approval gates, and
audit trails that keep every model compliant on its way into production. You'll
also own the trust side of the system: building the safety guardrails that keep
model outputs safe from misuse and treating user privacy as a design constraint
rather than an afterthought. As a senior member of the team, you'll mentor other
engineers and help set the technical standards the rest of the team builds
against. This is a role for someone who's comfortable operating at the
intersection of ML and distributed systems, as much at home tuning GPU
utilization and KV-cache for low-latency inference as designing the versioning
strategy that makes a rollback safe.
MINIMUM QUALIFICATIONS
Master's degree in Computer Science, Engineering, or a related field 8+ years of
experience in Machine learning and software engineering Proven track record of
shipping production-grade ML/LLM systems Strong understanding of LLMs,
fine-tuning, prompt engineering, and RAG patterns Experience building pipelines
that process multimodal data (structured and image) and integrate ML model
inference, including LLMs and embedding models, for data enrichment and
transformation Hands-on experience deploying, serving, and optimizing LLMs or ML
models in production, including inference runtimes/compilers (ONNX Runtime,
TensorRT/TensorRT-LLM), serving frameworks (Triton, vLLM, SGLang, TorchServe, or
similar), and tuning batching, KV-cache, and GPU utilization for low-latency,
high-throughput inference Experience with vector search technologies (e.g.,
Pinecone, Milvus) and storing/serving embeddings (e.g., pgvector, FAISS)
Experience with feature stores (e.g., Feast) and data lineage tracking Strong
proficiency in Python, with solid software engineering fundamentals, including
backend service frameworks (e.g., Flask, FastAPI), for building ML/LLM services,
pipelines, and tooling Working proficiency in Java or Scala, sufficient to
integrate with JVM-based data infrastructure (e.g., Spark, Flink, Kafka clients)
and the broader services platform. Experience with distributed systems, cloud
platforms (e.g., AWS), container orchestration (Kubernetes), CI/CD pipelines,
and building Data Pipelines on Spark using Airflow Experience with ML lifecycle
management and versioning practices, including experiment tracking, model
registry, deployment automation, and dataset/model versioning tools (e.g., DVC,
MLflow, Weights & Biases, Delta Lake) Experience with workflow orchestration
platforms (Airflow) Excellent communication skills and a collaborative,
team-oriented mindset
PREFERRED QUALIFICATIONS
Ph.D. in Computer Science, Machine Learning, or a related field Experience with
Go Solid understanding of machine learning algorithms, model evaluation metrics,
and data processing pipelines Active participation in open-source projects
related to AI/ML or backend development Familiarity with graph databases such as
TigerGraph Experience defining SLAs, quality metrics, and observability
standards for large-scale data platforms, with hands-on use of
monitoring/alerting tooling (e.g., Prometheus/Grafana, Datadog, or
OpenTelemetry-based tracing). Track record of mentoring engineers and
influencing technical direction across a team or organization Working knowledge
of data privacy principles and practices (e.g., data minimization, access
controls, privacy-preserving measurement) and experience applying them to ML
data pipelines Experience implementing model governance frameworks, including
approval workflows, audit trails, and compliance controls Experience
implementing safety guardrails for LLM-powered systems, including content
moderation, prompt-injection defenses, and red-teaming or adversarial evaluation
practices Hands-on experience with observability and evaluation tools for LLMs
(e.g., LangSmith, Weights & Biases, MLflow)
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Machine learning jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Gen AI Engineer -Dallas, TX
Photon · Dallas, Texas, United States
Automation & AI Engineer
ECS Tech Inc · Fairfax, Virginia, United States
AI Engineer 5
Capital One · San Jose, California, United States
Senior Applied AI Engineer
QuEra Computing Inc. · Boston, Massachusetts, United States
AI Engineer 5 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning)
Capital One · San Jose, California, United States
Senior Context Fusion AI Engineer - Autonomous Vehicles
NVIDIA · Redmond, Nevada, United States
Role information can change. Confirm current details on the original application page.
