Prepin
Log in
Apple

engineering opportunity

Machine Learning Engineer, Foundation Model Services

You will build production-grade solutions to serve foundation models in real-time and collaborate with researchers to develop inference for cutting-edge architectures. Additionally, you will create tools to identify and resolve performance bottlenecks across hardware and various use cases.

Seattle, Washington, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

Do you think differently? Are you eager to break the status quo, bold and

ambitious, unafraid to take risks, and passionate about building best-in-class

technology? If so, there's no better place to do it than Apple. The Foundation

Model Services team builds the frameworks, services, and tools that run Apple's

largest foundation models in production. Our infrastructure powers intelligent

experiences across products people use every day — Search, Music, TV, the App

Store, Messages, Photos, Spotlight, Safari, Siri, and more — serving millions of

queries at incredibly low latency while drawing every ounce of performance from

our hardware. Join us and you'll help bring intelligence to billions of users

around the world, working on optimizing and serving large language, vision, and

speech models at Apple's scale.

DESCRIPTION

Work closely with product teams to build production-grade solutions that launch

models serving customers in real time. Partner with foundation model researchers

to prototype and develop inference for cutting-edge model architectures, and

build tools that help us understand and remove performance bottlenecks across

different hardware and use cases. Write high-quality code, learn quickly in a

fast-moving field, and grow your impact as you take on larger pieces of the

system.

MINIMUM QUALIFICATIONS

2+ years of industry experience building and shipping production software and/or

machine learning systems. Proficiency in a modern programming language such as

Go or Python. Experience deploying and operating services on a cloud platform

(AWS, Azure, GCP, or equivalent) using containers and Kubernetes/Docker. 5 year+

industry experience in ML technologies (LLMs, Machine Learning, NLP, Information

Retrieval, Statistics). Experience building or operating high-throughput,

low-latency services. Strong communication and collaboration skills, with the

ability to partner across research and product teams. Bachelor’s degree or

higher in Computer Science or related technical field.

PREFERRED QUALIFICATIONS

Familiarity with Nvidia TensorRT-LLM, vLLLM, DeepSpeed, Nvidia Triton Server

etc.

Which skills does this role require?

Machine LearningPythonGoKubernetesDockerCloud PlatformsLLMsNLPInformation RetrievalStatisticsvLLMInference OptimizationMachine Learning EngineerFoundation ModelsInferenceProduction-gradeAWSAzureGCPContainersPerformance BottlenecksHigh-throughputLow-latencyComputer SciencePrototyping

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.