About the role
What will you do at Decart?
About the role
You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes.
We're looking for engineers with a demonstrated record in large-scale systems engineering and low-level optimization.
Minimum requirements Bachelor's degree in Electrical/Computer Engineering, Computer Science, or a related field, plus 2+ years of relevant experience (or equivalent practical experience) 2+ years developing low-level software in C/C++ (Python proficiency a plus) Solid grounding in operating systems fundamentals (process/thread scheduling, synchronization, virtual memory), CPU/GPU architecture, and hardware/software co-design Working knowledge of PyTorch and machine learning algorithms, with a focus on engineering application Demonstrated ability to profile compute and memory behavior, diagnose bottlenecks, and validate improvements with rigorous measurement What we're looking for Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators You've authored kernels for an ML accelerator — not just consumed them Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks) Experience scaling training or inference workloads across multi-accelerator clusters You've read — or patched — the internals of an ML framework Projects you might work on Cut milliseconds off end-to-end token latency in DOS by restructuring attention and sampling paths for a new accelerator generation Design communication schedules that overlap compute with network transfer across multi-chip topologies Build analytical performance models to predict where the next 2x is hiding before writing a line of kernel code Trace a throughput regression from a framework-level symptom down to instruction scheduling in generated assembly — and fix it Port DOS's kernel suite to new hardware and close the gap to theoretical peak
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
AI Workflow Engineer
Scout Motors Inc. · Charlotte, North Carolina, United States
AI Evaluations Engineer, US Decision Intelligence
Apple · Cupertino, California, United States
Research Engineer, Responsible Frontier AI Research, DeepMind
Google · New York, New York, United States
AI Solutions Engineer
Superior Essex · Sandy Springs, Georgia, United States
Applied AI Design Engineer (100 % remote) (m/f/d)
EWOR · Capon Bridge, West Virginia, United States
AI Outcome Customer Engineer, Forward Deployed Engineering
Google · Atlanta, Georgia, United States
Role information can change. Prepin can help you prepare, but does not submit an application for this role.