Prepin
Log in
Apple

engineering opportunity

AI Data Platform Engineer

Design, build, and maintain scalable AI data platforms and pipelines to support GenAI, Agentic AI, and Embodied AI solutions. Collaborate with cross-functional teams to develop data quality frameworks, metadata management, and production-ready AI datasets.

Cupertino, California, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Apple?

Imagine what you could do here. At Apple, we believe new insights have a way of

becoming excellent products, services, and customer experiences very quickly.

Bring passion and dedication to your job and there’s no telling what you could

accomplish. The people here at Apple don’t just build products — they build the

kind of wonder that’s revolutionized entire industries. It’s the diversity of

those people and their ideas that inspires the innovation that runs through

everything we do, from amazing technology to industry-leading environmental

efforts. Join Apple, and help us leave the world better than we found it.

Manufacturing Systems and Infrastructure (MSI) team is an engineering

organization under the Product Operations org. MSI is responsible for the

design, development, and maintenance of systems tools, services, and

applications required to efficiently run manufacturing operations at scale

across global factory sites. As an AI Data Platform Engineer with the MSI team,

you will design, build, and operate scalable AI data platforms that enable

GenAI, Agentic AI, and Embodied AI solutions across the enterprise. You will

develop reusable platform services, data pipelines, and data quality frameworks

that transform fragmented enterprise and multimodal data into trusted, AI-ready

datasets — combining expertise in AI data platform engineering, data quality,

systems engineering, and AI data lifecycle management to accelerate AI

innovation.

DESCRIPTION

Design, build, and maintain scalable AI data platforms, services, and APIs that

support and enable AI model development and production. Develop data ingestion,

transformation, and publishing pipelines for structured, unstructured, and

multimodal data. Build AI-ready datasets through ground truth creation, data

curation, annotation workflows, dataset versioning, and metadata management.

Develop data quality frameworks, validation pipelines, observability, and

evaluation metrics to ensure trusted AI datasets. Design and implement

Retrieval-Augmented Generation (RAG) pipelines, embedding workflows, vector

database integrations, and metadata services for enterprise AI applications.

Build scalable platform capabilities for managing the end-to-end AI data

lifecycle, including ground truth dataset creation, dataset versioning, metadata

and lineage management, automated data quality validation, governance, and

secure publishing of AI-ready datasets. Collaborate with AI/ML engineers,

software engineers, product teams, and domain experts to define AI data

requirements and deliver production-ready data solutions. Optimize platform

scalability, reliability, performance, security, and cost across cloud-native

environments. Drive engineering best practices for AI data architecture,

platform design, automation, testing, monitoring, and operational excellence.

Evaluate emerging AI technologies and continuously improve platform capabilities

that enable GenAI, agentic AI, and embodied AI solutions.

MINIMUM QUALIFICATIONS

Bachelor's or Master's degree in Computer Science, Software Engineering, Data

Engineering, or a related field. 5+ Experience designing and building scalable

data platforms and distributed systems. Strong programming skills in Python and

SQL, with proficiency in Java or Scala preferred. Experience with Airflow,

Kubeflow, or MLflow to build and orchestrate scalable AI data pipelines.

Experience building scalable batch and streaming data pipelines using Spark

(PySpark), Kafka, Airflow, and Ray, with proficiency in Pandas and modern data

lake/lakehouse architectures (e.g., Iceberg, Delta Lake). Hands-on experience

with AI data engineering, including ground truth dataset creation, data

curation, annotation pipelines, dataset versioning, and metadata management.

Experience implementing data validation, quality frameworks, observability, and

AI dataset evaluation. Knowledge of RAG architectures, embedding generation,

vector databases, and AI data preparation for LLMs and agentic AI. Experience

with cloud platforms (AWS, Azure, or GCP), Kubernetes, Docker, CI/CD, and

Infrastructure as Code. Strong understanding of distributed systems, APIs,

microservices, and enterprise integration patterns. Excellent communication,

collaboration, and technical leadership skills.

PREFERRED QUALIFICATIONS

Experience building platforms supporting GenAI, Agentic AI, or Embodied AI

applications. Experience with multimodal datasets, knowledge graphs, AI

evaluation frameworks, or vector search technologies. Familiarity with

enterprise data governance, lineage, metadata management, and AI compliance.

Experience working with manufacturing, operational, IoT, or industrial data

platforms. Demonstrated ability to lead technical initiatives and mentor

engineers.

Which skills does this role require?

PythonSQLData Quality FrameworksAI Data PlatformData QualityRetrieval-Augmented GenerationVector DatabaseJavaScalaCloud-nativeMultimodal DataManufacturing SystemsMachine Learning

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.