About the role
What will you do at Apple?
Imagine what you could do here. At Apple, great ideas have a way of becoming
extraordinary products, services, and customer experiences very quickly. Bring
passion and dedication to your work, and there’s no telling what you could
accomplish. The Channel Sales AI Product Engineering team is looking for a
Machine Learning Evaluation Engineer to help build and scale evaluation
capabilities for our next generation of AI-powered experiences. In this role,
you will develop evaluation frameworks, datasets, tooling, and quality signals
that enable teams to understand and continuously improve Generative AI and
LLM-powered products. You will work closely with Machine Learning, Software
Engineering, Quality Engineering, Product, Human Interface, Data Science, and
domain experts to establish rigorous evaluation practices throughout the AI
product lifecycle. You will help define how we measure the quality of AI
experiences across the Commerce domain, including Store AI, Shopping AI,
Learning AI, Content GenAI, Conversational AI, and Platform Self-Service. This
is an opportunity to work at the intersection of machine learning, software
engineering, data, and product quality, helping ensure our AI experiences are
accurate, relevant, grounded, reliable, and useful for users around the world.
DESCRIPTION
As a Machine Learning Evaluation Engineer, you will design and build scalable
evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI
products. You will: ◦ Design and develop automated evaluation frameworks and
pipelines for AI-powered products. ◦ Define evaluation methodologies and quality
metrics across dimensions such as accuracy, relevance, groundedness,
completeness, consistency, instruction following, and task completion. ◦ Build
and maintain high-quality evaluation datasets, including golden datasets,
benchmark sets, regression suites, adversarial scenarios, and production-derived
test sets. ◦ Develop Auto Eval capabilities that enable teams to rapidly
evaluate models, prompts, retrieval systems, agents, and end-to-end AI
experiences. ◦ Design and implement model-based evaluation approaches, including
LLM-as-a-Judge, while developing appropriate calibration and validation
methodologies. ◦ Develop Human-in-the-Loop (HITL) evaluation approaches for
complex or subjective quality dimensions where automated evaluation alone is
insufficient. ◦ Define evaluation rubrics, annotation guidelines, grading
criteria, and quality standards in partnership with product teams, domain
experts, and annotation teams. ◦ Build mechanisms to calibrate automated
evaluators against human judgment and measure evaluator consistency and
reliability. ◦ Evaluate end-to-end AI systems, including retrieval, context
construction, prompts, model responses, tool use, APIs, and downstream product
experiences. ◦ Develop evaluation methodologies for multi-turn conversations,
personalization, recommendations, tool use, reasoning, and agentic task
execution. ◦ Perform detailed error analysis and failure-mode investigation to
identify opportunities for model, prompt, retrieval, dataset, and product
improvements. ◦ Build reusable evaluation infrastructure, APIs, dashboards, and
developer tooling that can scale across multiple AI products and teams. ◦
Integrate evaluation into development and CI/CD workflows, enabling automated
regression detection, quality gates, and release-readiness assessments. ◦
Connect offline evaluation results with production signals to continuously
improve evaluation coverage and product quality. ◦ Partner closely with Machine
Learning, Software Engineering, Product, Quality Engineering, Human Interface,
and Data Science teams throughout research, development, evaluation, launch, and
continuous improvement.
MINIMUM QUALIFICATIONS
Typically requires a minimum of 7 years of related experience in Machine
Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality
Engineering, or a related technical field. Strong programming skills in Python
and experience developing production-quality software, ML systems, data
pipelines, or evaluation infrastructure. Experience developing or evaluating
LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other
machine-learning-driven products. Experience designing automated ML evaluation
frameworks, metrics, benchmarks, datasets, or experimentation methodologies.
Understanding of modern LLM application architectures, including prompting,
embeddings, retrieval-augmented generation (RAG), tool use, and agentic
workflows. Experience with model-based evaluation techniques and an
understanding of the strengths and limitations of approaches such as
LLM-as-a-Judge. Experience with Human-in-the-Loop evaluation, annotation, or
data-quality workflows. Strong understanding of statistical analysis,
experimentation, sampling, and measurement methodologies. Experience performing
model error analysis, failure analysis, and root-cause investigation. Ability to
work effectively across Machine Learning, Engineering, Product, Quality, and
Data teams. Excellent written and verbal communication skills, with the ability
to translate complex technical findings into clear, actionable recommendations.
Bachelor's degree in Computer Science, Machine Learning, Artificial
Intelligence, Data Science, Statistics, Electrical Engineering, or a related
technical field, or equivalent industry experience.
PREFERRED QUALIFICATIONS
Experience building evaluation infrastructure for production-scale LLM or
Generative AI applications. Experience evaluating RAG, conversational systems,
AI agents, personalization, recommendations, or multimodal AI. Experience
building golden datasets, regression suites, automated quality gates, and
continuous evaluation pipelines. Experience integrating ML evaluation into CI/CD
and production release processes. Experience with prompt evaluation, model
comparison, experiment tracking, and AI observability. Experience evaluating
multilingual AI experiences across languages, locales, and markets. Familiarity
with responsible AI evaluation, including robustness, safety, bias, and
adversarial testing. Experience developing internal ML platforms, developer
tooling, or self-service evaluation capabilities used across multiple teams.
Experience working with large-scale datasets and distributed ML or
data-processing infrastructure. Master's degree in Computer Science, Machine
Learning, Artificial Intelligence, Data Science, Statistics, Electrical
Engineering, or a related technical field, or equivalent industry experience.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Machine learning jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Gen AI Engineer -Dallas, TX
Photon · Dallas, Texas, United States
Senior Applied AI Engineer
QuEra Computing Inc. · Boston, Massachusetts, United States
GTM AI Engineer -Deal Desk
Motive Agency · United States
AI Engineer 5 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning)
Capital One · San Jose, California, United States
AI Engineer 5
Capital One · San Jose, California, United States
AI Engineer 3 (AI Foundations)
Capital One · San Jose, California, United States
Role information can change. Confirm current details on the original application page.
