Prepin
Log in
Apple

engineering opportunity

Sr. Machine Learning Engineer, Speech LLM Evaluation

You will own the data and metrics foundation for evaluating speech LLMs across accuracy, robustness, and conversational quality. This involves building evaluation datasets and designing automated judges to ensure models meet rigorous standards before deployment.

Cupertino, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

Join the team redefining what a deeply personal and integrated assistant can

be. As part of the Siri organization, you will help shape one of the world's

most widely used AI assistants, powered by our next-generation of Apple

Intelligence, with capabilities like personal context understanding and

on-screen awareness, built with privacy from the ground up. Your work will have

direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and

visionOS. Our Speech Evaluation team sits at the center of Apple's ASR, TTS, and

real-time conversational AI efforts, partnering directly with the modeling

teams. We're growing the team to take on a role focused specifically on

evaluating audio LLMs: designing the datasets that stress-test them and the

metrics that decide whether they're ready. You'll help define how Apple measures

a new class of models that listen, speak, and reason. This is a rare opportunity

to build at the intersection of cutting-edge AI and human-centered design,

shipping technology that is centered around users and their needs.

DESCRIPTION

This role owns the data and metrics foundation for evaluating speech LLMs (e.g.,

real-time speech understanding and generation models) across accuracy,

robustness, and conversational quality. You'll build and curate evaluation

datasets that reflect real usage — from personalized named-entity queries to

multi-turn fluid conversations — and design the metrics and automated judges

that turn model outputs into actionable, trustworthy signal. You'll work closely

with modeling, infrastructure, and product partners to make sure every new model

is evaluated quickly, consistently, and at the right level of rigor before it

reaches customers.

MINIMUM QUALIFICATIONS

Bachelor's degree in Computer Science, Electrical Engineering, or a related

field, or equivalent practical experience. Experience building or working with

text, speech or audio evaluation pipelines and metrics. Proficiency in Python

and experience building data processing pipelines at scale. Experience curating

or annotating datasets for machine learning evaluation or training. Working

knowledge of statistics as applied to measuring model performance and

interpreting evaluation results. Familiarity with large language model

evaluation techniques, including automated (LLM-as-judge) and human evaluation

methods. Strong written and verbal communication skills, with the ability to

explain evaluation results to both technical and non-technical audiences.

PREFERRED QUALIFICATIONS

Experience evaluating audio-native or multimodal (speech-in, speech-out) large

language models. Experience designing or running human evaluation studies (e.g.,

side-by-side comparisons, MOS ratings) at scale. Familiarity with

personalization and named-entity evaluation challenges in speech systems.

Experience with multilingual or international audio dataset development.

Experience with distributed data processing frameworks (e.g., Spark) for

large-scale audio dataset generation. Publication record or demonstrated

contributions in speech, audio ML, or NLP evaluation.

Which skills does this role require?

Machine LearningSpeech RecognitionLLM EvaluationPythonStatisticsASRTTSConversational AIData CurationModel Performance MeasurementMultimodal ModelsSpeech LLMData PipelinesLLM-as-judgeNamed-entity RecognitionApple IntelligenceModel EvaluationRobustnessAccuracyAudio ProcessingLLMs

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.