About the role
What will you do at Apple?
Join the team redefining what a deeply personal and integrated assistant can
be. As part of the Siri organization, you will help shape one of the world's
most widely used AI assistants, powered by our next-generation of Apple
Intelligence, with capabilities like personal context understanding and
on-screen awareness, built with privacy from the ground up. Your work will have
direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and
visionOS. Our Speech Evaluation team sits at the center of Apple's ASR, TTS, and
real-time conversational AI efforts, partnering directly with the modeling
teams. We're growing the team to take on a role focused specifically on
evaluating audio LLMs: designing the datasets that stress-test them and the
metrics that decide whether they're ready. You'll help define how Apple measures
a new class of models that listen, speak, and reason. This is a rare opportunity
to build at the intersection of cutting-edge AI and human-centered design,
shipping technology that is centered around users and their needs.
DESCRIPTION
This role owns the data and metrics foundation for evaluating speech LLMs (e.g.,
real-time speech understanding and generation models) across accuracy,
robustness, and conversational quality. You'll build and curate evaluation
datasets that reflect real usage — from personalized named-entity queries to
multi-turn fluid conversations — and design the metrics and automated judges
that turn model outputs into actionable, trustworthy signal. You'll work closely
with modeling, infrastructure, and product partners to make sure every new model
is evaluated quickly, consistently, and at the right level of rigor before it
reaches customers.
MINIMUM QUALIFICATIONS
Bachelor's degree in Computer Science, Electrical Engineering, or a related
field, or equivalent practical experience. Experience building or working with
text, speech or audio evaluation pipelines and metrics. Proficiency in Python
and experience building data processing pipelines at scale. Experience curating
or annotating datasets for machine learning evaluation or training. Working
knowledge of statistics as applied to measuring model performance and
interpreting evaluation results. Familiarity with large language model
evaluation techniques, including automated (LLM-as-judge) and human evaluation
methods. Strong written and verbal communication skills, with the ability to
explain evaluation results to both technical and non-technical audiences.
PREFERRED QUALIFICATIONS
Experience evaluating audio-native or multimodal (speech-in, speech-out) large
language models. Experience designing or running human evaluation studies (e.g.,
side-by-side comparisons, MOS ratings) at scale. Familiarity with
personalization and named-entity evaluation challenges in speech systems.
Experience with multilingual or international audio dataset development.
Experience with distributed data processing frameworks (e.g., Spark) for
large-scale audio dataset generation. Publication record or demonstrated
contributions in speech, audio ML, or NLP evaluation.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Machine learning jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Gen AI Engineer -Dallas, TX
Photon · Dallas, Texas, United States
Senior Applied AI Engineer
QuEra Computing Inc. · Boston, Massachusetts, United States
Automation & AI Engineer
ECS Tech Inc · Fairfax, Virginia, United States
AI Engineer 5 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning)
Capital One · San Jose, California, United States
AI Engineer 5
Capital One · San Jose, California, United States
Senior Context Fusion AI Engineer - Autonomous Vehicles
NVIDIA · Redmond, Nevada, United States
Role information can change. Confirm current details on the original application page.
