About the role
What will you do at Apple?
The On-Device Machine Learning team at Apple is responsible for enabling the
Research to Production lifecycle of cutting edge machine learning models that
power magical user experiences on Apple’s hardware and software platforms. Apple
is the best place to do on-device machine learning, and this team sits at the
heart of that discipline, interfacing with research, SW engineering, HW
engineering, and products. The On-device ML Performance team has the
responsibility to analyze latency, memory, power and numerical correctness of
the latest machine learning models running on Apple SoC’s, and to make Apple’s
ML software stack take full advantage of the capabilities in Apple’s ML
accelerators. The work from this cross functional team enables model developers’
decisions to optimize performance via advanced techniques such as quantization,
sparsity, performance and accuracy tradeoffs. The work of this team impacts all
new Apple HW and ML Inference on them. Our group is looking for an On-device ML
Performance Engineer, with technical expertise in computer architecture,
performance, memory, power, ML model architectures and on-device ML inference.
The role entails deep analysis of ML inference from the SW stack and low level
drivers to HW debug involving CPU, GPU, Apple Neural Engine, system memory and
power.
DESCRIPTION
As an engineer in this role, you will be primarily focused on analyzing and
optimizing the performance of the latest ML models on the latest iPhones and
Mac’s. You will work with models created by the most popular ML frameworks
(PyTorch, MLX, etc) and will analyze the inference of those models on device to
ensure the stack achieves full machine performance on Apple Silicon. The role
also includes scripting, coding, and generation of utilities and debug tools to
extract, analyze, and report performance and power related metrics for Apple HW.
The ideal candidate will have a passion for ML model architectures and ML
inference, deep knowledge of GPU and CPU, computer architecture and memory,
compilers and HW drivers.
MINIMUM QUALIFICATIONS
Experience with ML inference, quantization, performance and accuracy Familiarity
and experience with the most popular ML architectures (e.g. LLM’s, Diffusion
models, CNN’s) A passion to explore and learn about the latest advances in ML
model design and architecture, particularly as related to model implementation
on HW and on-device inference Familiarity with Operating Systems, embedded
systems, and CPU/GPU HW architectures Highly proficient in Python/C++ and shell
scripting Familiarity with Linux or macOS Exceptional clarity in verbal and
written communication, including the ability to present and lead discussions in
larger groups
PREFERRED QUALIFICATIONS
Masters or PhDs in Computer Science or relevant disciplines. Experience with
Apple’s CoreML, MPS Graph, Metal Performance Shader’s or MLX frameworks
Experience with any on-device ML stack, such as TFLite, ONNX, ExecuTorch, etc.
Experience with any ML authoring framework (PyTorch, TensorFlow, JAX, etc.).
Experience with Apple’s App development framework such as Xcode, Swift,
Objective-C Experience with any compiler stack (MLIR/LLVM/TVM etc.)
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Machine learning jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior Applied AI Engineer
QuEra Computing Inc. · Boston, Massachusetts, United States
Automation & AI Engineer
ECS Tech Inc · Fairfax, Virginia, United States
Senior Context Fusion AI Engineer - Autonomous Vehicles
NVIDIA · Redmond, Nevada, United States
Gen AI Engineer -Dallas, TX
Photon · Dallas, Texas, United States
AI Engineer 5 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning)
Capital One · San Jose, California, United States
AI Engineer 5
Capital One · San Jose, California, United States
Role information can change. Confirm current details on the original application page.
