Prepin
Log in
Apple

engineering opportunity

Staff/Sr. Machine Learning Engineer, Foundation Models - AI, Search & Knowledge Platforms

You will work alongside the Foundation Model Research team to optimize inference for cutting-edge model architectures and build production-grade solutions. Additionally, you will develop tools to identify performance bottlenecks and mentor other engineers within the organization.

Santa Clara, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

We are Foundation Model Inference Team, within AI, Search & Knowledge Platform

Technologies organization. Our team is responsible to build Inference stack to

power Apple Intelligence. It builds frameworks, services and tools that power

the largest Apple foundation models on servers. Our Infrastructure powers a wide

gamut of services at Apple including Apple Search, Apple Music, AppleTV,

AppStore, iMessages, Photos & Camera, Spotlight, Safari, Siri and upcoming ever

exciting Apple products serving millions of queries every day with incredible

low latencies, drawing every ounce of compute from our hardware. As part of this

group, you will get a chance to bring Intelligence to billions of users across

the world. You will have an opportunity to make difference in life of people by

empowering them with AI. You will have a chance to work on optimizing billions

of parameter language and vision and speech models using state of the art

technologies and make it run at scale of Apple.

DESCRIPTION

Work along side Foundation Model Research team to optimize inference for cutting

edge model architectures. Work closely with product teams to build Production

grade solutions to launch models serving millions of customers in real time.

Build tools to understand bottlenecks in Inference for different hardwares and

use cases. Mentor and guide engineers in the organization.

MINIMUM QUALIFICATIONS

5+ years of experience leading and driving complex, ambiguous projects.

Experience with LLM inference stack Familiarity with GPU programming concepts

using CUDA. Familiarity with one of the popular ML Frameworks like Pytorch,

Tensorflow. Have experience with high throughput services particularly at

supercomputing scale. Proficient with running applications on Cloud (AWS / Azure

or equivalent) using Kubernetes, Docker etc. Familiar with one of the popular ML

Frameworks like Pytorch, Tensorflow. BS in Computer Science, Artificial

Intelligence, Machine Learning, Information Retrieval, Data Science or related

field

PREFERRED QUALIFICATIONS

Proficient in building and maintaining systems written in modern languages (eg:

Golang, Python) Familiar with fundamental Deep Learning architectures such as

Transformers, Encoder/Decoder models. Familiarity with Nvidia TensorRT-LLM,

vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA

kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial

Intelligence, Machine Learning, Information Retrieval, Data Science or related

field.

Which skills does this role require?

LLM InferenceGPU ProgrammingPytorchTensorflowKubernetesDockerFoundation ModelsInferenceGPUCloud ComputingAWSAzureArtificial IntelligenceScalabilityHigh ThroughputOptimizationSoftware EngineeringInfrastructureGoPyTorchTensorFlowLLMs

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.