About the role
What will you do at Apple?
We are Foundation Model Inference Team, within AI, Search & Knowledge Platform
Technologies organization. Our team is responsible to build Inference stack to
power Apple Intelligence. It builds frameworks, services and tools that power
the largest Apple foundation models on servers. Our Infrastructure powers a wide
gamut of services at Apple including Apple Search, Apple Music, AppleTV,
AppStore, iMessages, Photos & Camera, Spotlight, Safari, Siri and upcoming ever
exciting Apple products serving millions of queries every day with incredible
low latencies, drawing every ounce of compute from our hardware. As part of this
group, you will get a chance to bring Intelligence to billions of users across
the world. You will have an opportunity to make difference in life of people by
empowering them with AI. You will have a chance to work on optimizing billions
of parameter language and vision and speech models using state of the art
technologies and make it run at scale of Apple.
DESCRIPTION
Work along side Foundation Model Research team to optimize inference for cutting
edge model architectures. Work closely with product teams to build Production
grade solutions to launch models serving millions of customers in real time.
Build tools to understand bottlenecks in Inference for different hardwares and
use cases. Mentor and guide engineers in the organization.
MINIMUM QUALIFICATIONS
5+ years of experience leading and driving complex, ambiguous projects.
Experience with LLM inference stack Familiarity with GPU programming concepts
using CUDA. Familiarity with one of the popular ML Frameworks like Pytorch,
Tensorflow. Have experience with high throughput services particularly at
supercomputing scale. Proficient with running applications on Cloud (AWS / Azure
or equivalent) using Kubernetes, Docker etc. Familiar with one of the popular ML
Frameworks like Pytorch, Tensorflow. BS in Computer Science, Artificial
Intelligence, Machine Learning, Information Retrieval, Data Science or related
field
PREFERRED QUALIFICATIONS
Proficient in building and maintaining systems written in modern languages (eg:
Golang, Python) Familiar with fundamental Deep Learning architectures such as
Transformers, Encoder/Decoder models. Familiarity with Nvidia TensorRT-LLM,
vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA
kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial
Intelligence, Machine Learning, Information Retrieval, Data Science or related
field.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Machine learning jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Staff AI Engineer - Personalization, Brand, Communications Tech
American Express · New York, New York, United States
AI Engineer 3 (AI Foundations)
Capital One · San Jose, California, United States
AI Engineer 5
Capital One · San Jose, California, United States
Lead AI Engineer
PepsiCo · Plano, Texas, United States
Gen AI Engineer -Dallas, TX
Photon · Dallas, Texas, United States
AI Engineer 5 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning)
Capital One · San Jose, California, United States
Role information can change. Confirm current details on the original application page.
