About the role
What will you do at Amazon?
Some CX shopping defects might be straightforward to detect and track. The
interesting ones are not because they depend on what a customer perceives. For
example, a search page may return legitimately different results, yet a shopper
has no way to tell apart. We turn these ambiguous perception questions into a
measurable artifact using LLMs, and we build the frameworks to know exactly
where model judgment can be trusted and where a human must decide and proving
it, against ground truth, at Amazon scale.
This is one example of an LLM based measurement pipeline you will own, but
that’s not all. You will extend that measurement to other parts of the shopping
experience like the homepage and the detail page, where the same customer
problem looks nothing like it does in search results, and where you will design
the measurement from scratch.
The larger goal is what makes this role unusual. Teams across Amazon are each
independently figuring out how to label quality with LLMs, hitting the same
problems alone: prompts that break on the next model version; golden sets nobody
audited, accuracy that collapses in other locales. Through the work above, you
will set the standard and build the production tooling behind it. Reusable
labeling pipelines, evaluation frameworks, and inference infrastructure that
hold up against Amazon-sized data and get adopted by teams who did not have to
use them.
Key job responsibilities
- Design and improve LLM-based labeling for perception-driven defects: prompt
design, sampling strategy, and the split between model judgment and human
annotation.
- Validate labeling quality against human ground truth, and build and maintain
the golden datasets that make that validation possible.
- Extend perceived-duplicate measurement to other parts of the shopping
experience, designing the methodology where none exists and evaluating
approaches already in use where one does.
- Build reusable, production-grade labeling and evaluation tooling, batch
inference, quality sampling, prompt and model version control, that operates on
Amazon-scale data.
- Define and publish the standards other teams adopt for using LLMs to measure
customer experience.
- Partner with science, engineering, and product teams across Stores to make
quality measurement usable in their decisions, and present metric results and
methodology changes to stakeholders.
A day in the life
We are a small team, which means your work is visible and your scope grows as
fast as you do. There are existing partnerships with other teams and a paved way
to cross org influence. Basic
Qualifications
- - PhD, or Master's degree and 4+
- years of CS, CE, ML or related field experience
- - Experience programming in Java, C++, Python or related language
- - Experience in any of the following areas: algorithms and data structures,
- parsing, numerical optimization, data mining, parallel and distributed
- computing, high-performance computing
Preferred Qualifications
- - Experience
- using Unix/Linux
- - Experience in professional software development
- Amazon is an equal opportunity employer and does not discriminate on the basis
- of protected veteran status, disability, or other legally protected status.
- Our inclusive culture empowers Amazonians to deliver the best results for our
- customers. If you have a disability and need a workplace accommodation or
- adjustment during the application and hiring process, including support for the
- interview or onboarding process, please visit
- https://amazon.jobs/content/en/how-we-hire/accommodations
- [https://amazon.jobs/content/en/how-we-hire/accommodations] for more
- information. If the country/region you’re applying in isn’t listed, please
- contact your Recruiting Partner.
- The base salary range for this position is listed below. Your Amazon package
- will include sign-on payments and restricted stock units (RSUs). Final
- compensation will be determined based on factors including experience,
- qualifications, and location. Amazon also offers comprehensive benefits
- including health insurance (medical, dental, vision, prescription, Basic Life &
- AD&D insurance and option for Supplemental life plans, EAP, Mental Health
- Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy
- Reimbursement coverage), 401(k) matching, paid time off, and parental leave.
- Learn more about our benefits at https://amazon.jobs/en/benefits
- [https://amazon.jobs/en/benefits].
- USA, WA, Seattle - 142,800.00 - 193,200.00 USD annually
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Data Scientist and AI Specialist
NexOne, Inc. · Clearfield, Utah, United States
AI Data Scientist II
Eagleview · United States
Foundational Model Research Data Scientist
Sapience AI · Seattle, Washington, United States
Data Scientist IV TS/SCI
NASA Jet Propulsion Laboratory · Pasadena, California, United States
Data Scientist IV TS/SCI
NASA Jet Propulsion Laboratory · United States
Principal Data Scientist, Director (New York)
Fitch Group, Inc. · New York, New York, United States
Role information can change. Confirm current details on the original application page.
