About the role
What will you do at Apple?
The Apple Intelligence Platform Experience Validation team builds the tooling
and automation that keeps Apple Intelligence features high-quality before they
ship. We are looking for a Senior SDET to lead the design and implementation of
automated model evaluation: standing up LLM-as-a-judge in existing and new
pipelines, and building the infrastructure that catches model regressions before
they reach human evaluation or the live on population. This is a hands-on,
senior individual-contributor role. You will own eval automation as a discipline
across the team, partnering with modeling, framework, and infrastructure teams
to make model quality a first-class, continuously measured signal.
DESCRIPTION
You will build and maintain model level, component or end-to-end evaluation
coverage for the generative features our team validates. Your job is to leverage
LLM judge scoring output quality in automation, ensuring reliable, repeatable
eval jobs that run that produce actionable signal. The kinds of problems you
will work on include: * Image / visual generation: validating model output and
its associated classification metadata, and detecting quality or behavior
regressions across model updates. * Natural-language generation: evaluating
whether generated artifacts and responses match user intent, moving at-desk LLM
judges into a scalable and repeatable automation environment. * Correctness
beyond string matching: replacing exact-match checks for open-ended or factual
responses with an LLM-as-judge stage integrated into the pipeline. * Generated
insights and summaries: assessing whether model-generated content is sensible
and good enough to surface to users. You will decide when a component-level
check (an API or CLI that exercises the model against its framework) is
sufficient and when a full end-to-end user flow is required, and you will build
the tooling for both.
MINIMUM QUALIFICATIONS
BS in Computer Science, Mathematics, or a related field (or equivalent practical
experience) Three years of relevant industry experience in test automation,
software development, or related areas.
PREFERRED QUALIFICATIONS
Strong practical knowledge of Python, including data-pipeline fluency
(JSON/YAML, REST APIs). Hands-on experience with LLM-as-a-judge evaluation and
rubric design, or a strong demonstrated ability to ramp into it quickly. Strong
software engineering fundamentals — able to define atomic, composable components
and build maintainable pipelines and tooling, not just scripts. Strong debugging
and triage skills; able to separate genuine regressions from infrastructure or
rubric noise. Strong knowledge of the software development lifecycle, testing
methodologies, and QA processes. Excellent written and verbal communication;
able to document clearly and describe quality signal to modeling and leadership
audiences. Ability to lead work across varying priorities and partner
multi-functionally with modeling, framework, and infrastructure teams.
Experience building on-device tooling and device/model eval infrastructure.
Experience integrating with CI/CD and job orchestration systems, and comfort
deploying tooling as reusable libraries. Familiarity with generative model
behavior — image generation, NLP, or LLM output evaluation. Experience curating
and reasoning about large datasets; comfort manually inspecting data (Jupyter or
similar) to build intuition and drive next steps. Awareness of dataset bias and
fairness considerations in evaluation. Experience with database/query tooling
(e.g., SQL) and dashboards/visualization for reporting quality trends.
Experience with Xcode is a bonus.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Machine learning jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Gen AI Engineer -Dallas, TX
Photon · Dallas, Texas, United States
GTM AI Engineer -Deal Desk
Motive Agency · United States
Automation & AI Engineer
ECS Tech Inc · Fairfax, Virginia, United States
Sr. Applied AI Engineer
phData · United States
Senior Applied AI Engineer
QuEra Computing Inc. · Boston, Massachusetts, United States
AI Engineer 5 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning)
Capital One · San Jose, California, United States
Role information can change. Confirm current details on the original application page.
