About the role
What will you do at Apple?
Do you have a passion for computer vision and solving deep learning problems?
The Video Engineering Data Analytics and Quality group is seeking an expert in
evaluating machine learning and deep learning models, including foundation
models and multimodal systems. This role will play a critical part in crafting
robust evaluation frameworks, using both traditional statistical methods and
modern techniques like LLM-as-a-Judge! The ideal candidate combines strong
analytical thinking, expertise in Python, and advanced knowledge of statistical
methodologies and data quality standards. This role involves collaboration with
teams at Apple passionate about developing foundation models, including ML
engineers, data scientists, and ML Infrastructure engineers to deliver amazing
user experiences!
DESCRIPTION
Develop robust methodologies to assess the performance of foundation models
(e.g., LLMs, vision-language models, etc.) across diverse tasks. Leverage LLMs
as judges to perform subjective and open-ended model evaluations (e.g., for
summarization, reasoning, or multimodal generation tasks). Build, curate, and
lead evaluation datasets and benchmarks. Advanced proficiency in at least one
scripting language, preferably Python. Collaborate with research, engineering,
and product teams to define evaluation goals aligned with user experience and
product quality. Conduct failure analysis and uncover edge cases to improve
model robustness. Contribute to our tools and infrastructure to automate and
scale evaluation processes.
MINIMUM QUALIFICATIONS
BS and a minimum of 3 years relevant industry experience Strong experience in
evaluating supervised, unsupervised, and deep learning models. Hands-on
experience evaluating LLMs and using them as scoring/judging mechanisms.
Familiarity with multimodal models (e.g., image + text, video + audio) and
related evaluation challenges. Proficiency in Python and libraries such as
NumPy, pandas, scikit-learn, PyTorch, or TensorFlow. Solid understanding of
statistical testing, sampling, confidence intervals, and metrics (e.g.,
precision/recall, BLEU, ROUGE, FID, etc.). Strong documentation skills,
including the ability to write technical reports and present to non-technical
audiences.
PREFERRED QUALIFICATIONS
Experience working with open-source evaluation tools like OpenEval, ELO-based
ranking, or LLM-as-a-Judge frameworks. Familiarity with prompt engineering,
few-shot or zero-shot evaluation techniques. Experience evaluating generative
models (e.g., text generation, image generation). Prior contributions to ML
benchmarks or public evaluations. Strong interpersonal skills.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Machine learning jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
AI and ML Engineer
Booz Allen Hamilton · Lorton, Virginia, United States
AI Engineer 3 (AI Foundations)
Capital One · San Jose, California, United States
Senior Context Fusion AI Engineer - Autonomous Vehicles
NVIDIA · Redmond, Nevada, United States
AI Engineer 5
Capital One · San Jose, California, United States
AI Engineer
Booz Allen Hamilton · Reston, Virginia, United States
Operational Technology AI Engineer
Booz Allen Hamilton · Chantilly, Virginia, United States
Role information can change. Confirm current details on the original application page.
