Prepin
Log in
Apple

engineering opportunity

Sr. Machine Learning Engineer, ML Systems Evaluation Engineering

You will lead the evaluation of Siri and AI/ML models at scale to improve user experiences while maintaining strict privacy standards. This involves collaborating with cross-functional teams to curate high-quality datasets and drive model development through rigorous offline evaluation insights.

Cupertino, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

Join the team redefining what a deeply personal and integrated assistant can

be. As part of the Siri organization, you will help shape one of the world's

most widely used AI assistants, powered by our next-generation of Apple

Intelligence, with capabilities like personal context understanding and

on-screen awareness, built with privacy from the ground up. Your work will have

direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and

visionOS. This is a rare opportunity to build at the intersection of

cutting-edge AI and human-centered design, shipping technology that is centered

around users and their needs. Join the SCALE (Statistical Coverage and

Large-scale Evaluation) team at Apple and contribute to a highly accomplished

group that evaluates Siri and AI/ML models at scale to delight and inspire our

users globally. If you are excited by large-scale systems, rigorous statistical

evaluation, natural language understanding, and shaping the future of AI, we

want you here! You’ll do more than join something, you’ll add something!

DESCRIPTION

We are seeking a highly skilled Senior Machine Learning Engineer specializing in

Conversational AI. Our goal is to deliver offline evaluation insights that drive

model development and improve the end-user experience, all while upholding

Apple's strict privacy standards. In this pivotal role, you will collaborate

with cross-functional teams to curate and evolve high-quality evaluation

datasets for state-of-the-art models. With the advent of Apple Intelligence, you

will tackle novel challenges in evaluating highly personalized user experiences.

MINIMUM QUALIFICATIONS

7+ years of professional experience applying machine learning to real-world

problems and crafting scalable data solutions, specifically in natural language

products. Proven experience managing large-scale datasets for ML training and/or

evaluation. Excellent programming skills in Python MS/PhD in Machine Learning,

Computer Science, or equivalent experience in a related field

PREFERRED QUALIFICATIONS

The qualifications that will benefit a candidate to be successful in your role.

Please list no more than 6-8. Deep domain knowledge in Conversational AI and a

strong understanding of the end-to-end ML product lifecycle. Expertise in

defining and measuring evaluation coverage for Large Language Models (LLMs) and

agentic systems. Track record of delivering large-scale, cross-functional ML

product or platform outcomes. Excellent problem-solving, critical thinking, and

communication skills to drive alignment across teams. Experience with systems

engineering; in-depth understanding of interdependencies of ML and SW components

Which skills does this role require?

Machine LearningPythonStatistical EvaluationData EngineeringSystems EngineeringNatural Language UnderstandingData CurationModel DevelopmentProblem SolvingCross-functional CollaborationSiriApple IntelligenceLLMPrivacySoftware EngineeringEvaluation MetricsModel LifecycleAI AssistantPersonalization

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.