Prepin
Log in
Apple

engineering opportunity

Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E

You will lead the design and implementation of automated model evaluation for Apple Intelligence features, building infrastructure to catch regressions. This involves creating end-to-end evaluation pipelines and leveraging LLM-as-a-judge to ensure high-quality model outputs.

Cupertino, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

The Apple Intelligence Platform Experience Validation team builds the tooling

and automation that keeps Apple Intelligence features high-quality before they

ship. We are looking for a Senior SDET to lead the design and implementation of

automated model evaluation: standing up LLM-as-a-judge in existing and new

pipelines, and building the infrastructure that catches model regressions before

they reach human evaluation or the live on population. This is a hands-on,

senior individual-contributor role. You will own eval automation as a discipline

across the team, partnering with modeling, framework, and infrastructure teams

to make model quality a first-class, continuously measured signal.

DESCRIPTION

You will build and maintain model level, component or end-to-end evaluation

coverage for the generative features our team validates. Your job is to leverage

LLM judge scoring output quality in automation, ensuring reliable, repeatable

eval jobs that run that produce actionable signal. The kinds of problems you

will work on include: * Image / visual generation: validating model output and

its associated classification metadata, and detecting quality or behavior

regressions across model updates. * Natural-language generation: evaluating

whether generated artifacts and responses match user intent, moving at-desk LLM

judges into a scalable and repeatable automation environment. * Correctness

beyond string matching: replacing exact-match checks for open-ended or factual

responses with an LLM-as-judge stage integrated into the pipeline. * Generated

insights and summaries: assessing whether model-generated content is sensible

and good enough to surface to users. You will decide when a component-level

check (an API or CLI that exercises the model against its framework) is

sufficient and when a full end-to-end user flow is required, and you will build

the tooling for both.

MINIMUM QUALIFICATIONS

BS in Computer Science, Mathematics, or a related field (or equivalent practical

experience) Three years of relevant industry experience in test automation,

software development, or related areas.

PREFERRED QUALIFICATIONS

Strong practical knowledge of Python, including data-pipeline fluency

(JSON/YAML, REST APIs). Hands-on experience with LLM-as-a-judge evaluation and

rubric design, or a strong demonstrated ability to ramp into it quickly. Strong

software engineering fundamentals — able to define atomic, composable components

and build maintainable pipelines and tooling, not just scripts. Strong debugging

and triage skills; able to separate genuine regressions from infrastructure or

rubric noise. Strong knowledge of the software development lifecycle, testing

methodologies, and QA processes. Excellent written and verbal communication;

able to document clearly and describe quality signal to modeling and leadership

audiences. Ability to lead work across varying priorities and partner

multi-functionally with modeling, framework, and infrastructure teams.

Experience building on-device tooling and device/model eval infrastructure.

Experience integrating with CI/CD and job orchestration systems, and comfort

deploying tooling as reusable libraries. Familiarity with generative model

behavior — image generation, NLP, or LLM output evaluation. Experience curating

and reasoning about large datasets; comfort manually inspecting data (Jupyter or

similar) to build intuition and drive next steps. Awareness of dataset bias and

fairness considerations in evaluation. Experience with database/query tooling

(e.g., SQL) and dashboards/visualization for reporting quality trends.

Experience with Xcode is a bonus.

Which skills does this role require?

Test AutomationLLM EvaluationData PipelinesSQLGenerative AIInfrastructure as CodeModel Regression TestingQuality AssuranceData AnalysisXcodeApple IntelligenceAutomationRegression TestingModel EvaluationFrameworksVisualizationDashboardingAPICLIMachine LearningLLMs

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.