Prepin
Log in
Apple

engineering opportunity

AI Engineer – Algorithm Evaluation & Agentic Systems

You will lead the benchmarking and integration of state-of-the-art models for image and video understanding within the DAQ team. Your role involves rigorously testing models in applied settings, uncovering edge-case failure modes, and architecting advanced agentic systems.

Sunnyvale, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

How do we ensure Apple's next-generation AI products are robust, safe, and truly

intelligent? Join the DAQ team to help answer that. We are seeking an AI

Engineer specializing in algorithm evaluation and agentic systems design for

advanced computer vision and video understanding algorithms. What We Value

Production mindset: correctness, observability and maintainability Ability to

reason about system-level tradeoffs, not just model performance Ability to

balance experimentation speed with engineering rigor Comfort working in

ambiguous problem spaces and defining metrics from first principles Clear

communication of technical findings to both technical and non-technical

audiences

DESCRIPTION

Within the DAQ team, our core mission is to evaluate and elevate advanced visual

technologies. As a key member of this group, you will lead the benchmarking and

integration of state-of-the-art models for image and video understanding. Rather

than focusing on core model training, you will apply your deep CV and ML

expertise to rigorously test models in applied settings, uncover edge-case

failure modes, and architect advanced agentic systems. If you are passionate

about AI safety, robust evaluation, and building autonomous multi-modal

workflows that bridge experimentation with production, we’d love to hear from

you.

MINIMUM QUALIFICATIONS

MS and a minimum of 3 years relevant industry experience 3+ years of applied

experience in Machine Learning, Computer Vision, or AI System Evaluation Solid

ML Foundation: Deep understanding of core Machine Learning principles, including

probability, statistics, data distributions, and model bias/variance. You can

apply statistical rigor to ensure evaluation metrics are meaningful and

reliable. Computer Vision Expertise: Deep theoretical and practical

understanding of Computer Vision (CV) and Vision-Language Models (VLMs). You

must understand how Vision Transformers (ViTs), spatial-temporal modeling, and

image/video processing work under the hood to effectively evaluate them.

Advanced Evaluation Skills: Proven track record of defining robust metrics/KPIs

and designing rigorous evaluation frameworks for generative AI or foundation

models. Deep experience with custom benchmark creation, automated regression

testing, LLM/VLM-as-a-judge methodologies, and human-in-the-loop evaluation.

Agentic Systems: Experience building and evaluating LLM/VLM-powered agents,

including tool use, multi-step reasoning, planning, and memory management

workflows. Failure Analysis: Strong intuition for probing ML models to discover

edge cases, hallucinations, and performance bottlenecks in constrained

environments. Be able to translate findings into actionable improvement

recommendations. Engineering Excellence: Strong proficiency in Python and

experience with deep learning frameworks (PyTorch) for running inference,

extracting embeddings, and building scalable evaluation pipelines.

PREFERRED QUALIFICATIONS

Demonstrated ability to lead technical evaluation strategies end-to-end, drive

architectural decisions for testing infrastructure, and mentor engineers. Strong

foundation in statistics, including hypothesis testing, confidence intervals,

and experimental design Knowledge of reinforcement learning, planning, or

decision-making systems Experience evaluating multi-modal or multi-agent systems

Prior work on AI reliability, safety, or benchmarking

Which skills does this role require?

Machine LearningComputer VisionVision-Language ModelsPythonPyTorchAlgorithm EvaluationAgentic SystemsStatistical AnalysisData DistributionsModel BiasRegression TestingHuman-in-the-loop EvaluationFailure AnalysisMulti-modal WorkflowsAI EngineerViTsLLMVLMInferenceEmbeddingsAI SafetyAutonomous WorkflowsKPIsObservabilityMaintainabilityLLMsA/B Testing

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.