Prepin
Log in
Decart

engineering opportunity

Kernel Engineer (ML Accelerators)

Kernel Engineer (ML Accelerators) role at Decart in San Francisco, US. Focus areas include C/C++, CUDA, C/C++, Python, PyTorch, CUDA. Solid grounding in operating systems fundamentals (process/thread scheduling, synchronization, virtual memory), CPU/GPU architecture, and hardware/software co-design Working knowledge of PyTorch and machine learning algorithms, with a focus on engineering...

San Francisco, USFull-time

Posted

About the role

What will you do at Decart?

About the role

You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes.

We're looking for engineers with a demonstrated record in large-scale systems engineering and low-level optimization.

Minimum requirements Bachelor's degree in Electrical/Computer Engineering, Computer Science, or a related field, plus 2+ years of relevant experience (or equivalent practical experience) 2+ years developing low-level software in C/C++ (Python proficiency a plus) Solid grounding in operating systems fundamentals (process/thread scheduling, synchronization, virtual memory), CPU/GPU architecture, and hardware/software co-design Working knowledge of PyTorch and machine learning algorithms, with a focus on engineering application Demonstrated ability to profile compute and memory behavior, diagnose bottlenecks, and validate improvements with rigorous measurement What we're looking for Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators You've authored kernels for an ML accelerator — not just consumed them Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks) Experience scaling training or inference workloads across multi-accelerator clusters You've read — or patched — the internals of an ML framework Projects you might work on Cut milliseconds off end-to-end token latency in DOS by restructuring attention and sampling paths for a new accelerator generation Design communication schedules that overlap compute with network transfer across multi-chip topologies Build analytical performance models to predict where the next 2x is hiding before writing a line of kernel code Trace a throughput regression from a framework-level symptom down to instruction scheduling in generated assembly — and fix it Port DOS's kernel suite to new hardware and close the gap to theoretical peak

Which skills does this role require?

C/C++CUDA

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Prepin can help you prepare, but does not submit an application for this role.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.