Prepin
Log in
Apple

engineering opportunity

Machine Learning Engineer - Agentic AI Evaluation Frameworks

You will design and build scalable evaluation frameworks, datasets, and quality signals to assess the performance of Generative AI and LLM-powered products. This involves collaborating with cross-functional teams to establish rigorous evaluation practices and improve the accuracy and reliability of AI experiences across the Commerce domain.

Cupertino, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

Imagine what you could do here. At Apple, great ideas have a way of becoming

extraordinary products, services, and customer experiences very quickly. Bring

passion and dedication to your work, and there’s no telling what you could

accomplish. The Channel Sales AI Product Engineering team is looking for a

Machine Learning Evaluation Engineer to help build and scale evaluation

capabilities for our next generation of AI-powered experiences. In this role,

you will develop evaluation frameworks, datasets, tooling, and quality signals

that enable teams to understand and continuously improve Generative AI and

LLM-powered products. You will work closely with Machine Learning, Software

Engineering, Quality Engineering, Product, Human Interface, Data Science, and

domain experts to establish rigorous evaluation practices throughout the AI

product lifecycle. You will help define how we measure the quality of AI

experiences across the Commerce domain, including Store AI, Shopping AI,

Learning AI, Content GenAI, Conversational AI, and Platform Self-Service. This

is an opportunity to work at the intersection of machine learning, software

engineering, data, and product quality, helping ensure our AI experiences are

accurate, relevant, grounded, reliable, and useful for users around the world.

DESCRIPTION

As a Machine Learning Evaluation Engineer, you will design and build scalable

evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI

products. You will: ◦ Design and develop automated evaluation frameworks and

pipelines for AI-powered products. ◦ Define evaluation methodologies and quality

metrics across dimensions such as accuracy, relevance, groundedness,

completeness, consistency, instruction following, and task completion. ◦ Build

and maintain high-quality evaluation datasets, including golden datasets,

benchmark sets, regression suites, adversarial scenarios, and production-derived

test sets. ◦ Develop Auto Eval capabilities that enable teams to rapidly

evaluate models, prompts, retrieval systems, agents, and end-to-end AI

experiences. ◦ Design and implement model-based evaluation approaches, including

LLM-as-a-Judge, while developing appropriate calibration and validation

methodologies. ◦ Develop Human-in-the-Loop (HITL) evaluation approaches for

complex or subjective quality dimensions where automated evaluation alone is

insufficient. ◦ Define evaluation rubrics, annotation guidelines, grading

criteria, and quality standards in partnership with product teams, domain

experts, and annotation teams. ◦ Build mechanisms to calibrate automated

evaluators against human judgment and measure evaluator consistency and

reliability. ◦ Evaluate end-to-end AI systems, including retrieval, context

construction, prompts, model responses, tool use, APIs, and downstream product

experiences. ◦ Develop evaluation methodologies for multi-turn conversations,

personalization, recommendations, tool use, reasoning, and agentic task

execution. ◦ Perform detailed error analysis and failure-mode investigation to

identify opportunities for model, prompt, retrieval, dataset, and product

improvements. ◦ Build reusable evaluation infrastructure, APIs, dashboards, and

developer tooling that can scale across multiple AI products and teams. ◦

Integrate evaluation into development and CI/CD workflows, enabling automated

regression detection, quality gates, and release-readiness assessments. ◦

Connect offline evaluation results with production signals to continuously

improve evaluation coverage and product quality. ◦ Partner closely with Machine

Learning, Software Engineering, Product, Quality Engineering, Human Interface,

and Data Science teams throughout research, development, evaluation, launch, and

continuous improvement.

MINIMUM QUALIFICATIONS

Typically requires a minimum of 7 years of related experience in Machine

Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality

Engineering, or a related technical field. Strong programming skills in Python

and experience developing production-quality software, ML systems, data

pipelines, or evaluation infrastructure. Experience developing or evaluating

LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other

machine-learning-driven products. Experience designing automated ML evaluation

frameworks, metrics, benchmarks, datasets, or experimentation methodologies.

Understanding of modern LLM application architectures, including prompting,

embeddings, retrieval-augmented generation (RAG), tool use, and agentic

workflows. Experience with model-based evaluation techniques and an

understanding of the strengths and limitations of approaches such as

LLM-as-a-Judge. Experience with Human-in-the-Loop evaluation, annotation, or

data-quality workflows. Strong understanding of statistical analysis,

experimentation, sampling, and measurement methodologies. Experience performing

model error analysis, failure analysis, and root-cause investigation. Ability to

work effectively across Machine Learning, Engineering, Product, Quality, and

Data teams. Excellent written and verbal communication skills, with the ability

to translate complex technical findings into clear, actionable recommendations.

Bachelor's degree in Computer Science, Machine Learning, Artificial

Intelligence, Data Science, Statistics, Electrical Engineering, or a related

technical field, or equivalent industry experience.

PREFERRED QUALIFICATIONS

Experience building evaluation infrastructure for production-scale LLM or

Generative AI applications. Experience evaluating RAG, conversational systems,

AI agents, personalization, recommendations, or multimodal AI. Experience

building golden datasets, regression suites, automated quality gates, and

continuous evaluation pipelines. Experience integrating ML evaluation into CI/CD

and production release processes. Experience with prompt evaluation, model

comparison, experiment tracking, and AI observability. Experience evaluating

multilingual AI experiences across languages, locales, and markets. Familiarity

with responsible AI evaluation, including robustness, safety, bias, and

adversarial testing. Experience developing internal ML platforms, developer

tooling, or self-service evaluation capabilities used across multiple teams.

Experience working with large-scale datasets and distributed ML or

data-processing infrastructure. Master's degree in Computer Science, Machine

Learning, Artificial Intelligence, Data Science, Statistics, Electrical

Engineering, or a related technical field, or equivalent industry experience.

Which skills does this role require?

Machine LearningPythonEvaluation FrameworksSoftware EngineeringNLPAgentic AIHuman-in-the-LoopStatistical AnalysisPrompt EngineeringError AnalysisModel EvaluationConversational AIModel-based EvaluationLLM-as-a-JudgeQuality EngineeringAPIDashboardsRegression TestingBenchmarkingRecommendation SystemsDistributed InfrastructureCommerce DomainLLMsA/B Testing

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.