Prepin
Log in
Apple

engineering opportunity

Machine Learning - Data Scientist

Develop and implement robust evaluation frameworks for foundation models, including LLMs and multimodal systems. Collaborate with cross-functional teams to define quality goals, conduct failure analysis, and automate evaluation processes.

Sunnyvale, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

Do you have a passion for computer vision and solving deep learning problems?

The Video Engineering Data Analytics and Quality group is seeking an expert in

evaluating machine learning and deep learning models, including foundation

models and multimodal systems. This role will play a critical part in crafting

robust evaluation frameworks, using both traditional statistical methods and

modern techniques like LLM-as-a-Judge! The ideal candidate combines strong

analytical thinking, expertise in Python, and advanced knowledge of statistical

methodologies and data quality standards. This role involves collaboration with

teams at Apple passionate about developing foundation models, including ML

engineers, data scientists, and ML Infrastructure engineers to deliver amazing

user experiences!

DESCRIPTION

Develop robust methodologies to assess the performance of foundation models

(e.g., LLMs, vision-language models, etc.) across diverse tasks. Leverage LLMs

as judges to perform subjective and open-ended model evaluations (e.g., for

summarization, reasoning, or multimodal generation tasks). Build, curate, and

lead evaluation datasets and benchmarks. Advanced proficiency in at least one

scripting language, preferably Python. Collaborate with research, engineering,

and product teams to define evaluation goals aligned with user experience and

product quality. Conduct failure analysis and uncover edge cases to improve

model robustness. Contribute to our tools and infrastructure to automate and

scale evaluation processes.

MINIMUM QUALIFICATIONS

BS and a minimum of 3 years relevant industry experience Strong experience in

evaluating supervised, unsupervised, and deep learning models. Hands-on

experience evaluating LLMs and using them as scoring/judging mechanisms.

Familiarity with multimodal models (e.g., image + text, video + audio) and

related evaluation challenges. Proficiency in Python and libraries such as

NumPy, pandas, scikit-learn, PyTorch, or TensorFlow. Solid understanding of

statistical testing, sampling, confidence intervals, and metrics (e.g.,

precision/recall, BLEU, ROUGE, FID, etc.). Strong documentation skills,

including the ability to write technical reports and present to non-technical

audiences.

PREFERRED QUALIFICATIONS

Experience working with open-source evaluation tools like OpenEval, ELO-based

ranking, or LLM-as-a-Judge frameworks. Familiarity with prompt engineering,

few-shot or zero-shot evaluation techniques. Experience evaluating generative

models (e.g., text generation, image generation). Prior contributions to ML

benchmarks or public evaluations. Strong interpersonal skills.

Which skills does this role require?

PythonMachine LearningDeep LearningComputer VisionStatistical MethodologiesData QualityNumPyPandasScikit-learnPyTorchTensorFlowMultimodal ModelsFailure AnalysisBenchmarkingData ScientistFoundation ModelsMultimodal SystemsELO-based RankingGenerative ModelsVideo EngineeringData AnalyticsSupervised LearningUnsupervised LearningPrecisionRecallBLEUROUGEFIDLLMs

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.