Prepin
Log in
Modular, a Qualcomm company

engineering opportunity

Senior AI Runtime Engineer

You will design and develop runtime and cross-stack optimizations to improve CPU and GPU efficiency for AI models. Additionally, you will collaborate with cross-functional teams to port the runtime stack to new hardware and ensure state-of-the-art performance.

United StatesremoteFULL_TIME

Posted

About the role

What will you do at Modular, a Qualcomm company?

ABOUT THE ROLE:

ML developers today face significant friction when deploying trained models.

They work in a fragmented space with incomplete, patchwork solutions that

require extensive performance tuning and model-specific optimizations. At

Modular, we are building the next-generation AI platform that will radically

improve how developers build and deploy AI models.

A core part of this offering is a platform that enables customers to achieve

state-of-the-art performance across model families and frameworks. As an AI

Runtime Engineer, you will own a runtime that operates on various CPU and GPU

hardware platforms, optimizing performance for diverse customer AI models.

LOCATION: Candidates based in the US or Canada are welcome to apply. You can

work in our office in Los Altos, CA or remotely from home. Onboarding for new

hires is conducted in-person in our Los Altos, CA office.

WHAT YOU WILL DO:

* Design and develop runtime and cross-stack optimizations to improve CPU and

GPU efficiency, addressing issues such as CPU overhead, caching, and data

locality across multiple devices.

* Port the Modular runtime stack to new hardware platforms and develop an API

to streamline this process.

* Collaborate with the compiler, kernels, serving, and models teams to design

core technologies that achieve state-of-the-art end-to-end performance on

various CPU and GPU hardware.

* Collaborate with the customer success team and engage with customers to

understand their performance requirements and use cases.

* Collaborate with tooling and infrastructure teams to design systems for

automated performance analysis and benchmarking.

WHAT YOU BRING TO THE TABLE:

* 5+ years of experience working on high-performance computing systems.

* Experience in C++ programming and complex software systems.

* Experience with CPU or GPU runtime optimizations and performance analysis on

CPUs, GPUs, or AI accelerators.

* Proficiency with one or more profiling tools (CPU or GPU).

* Creativity and curiosity for solving complex problems, a team-oriented

attitude that enables you to work well with others, and alignment with our

culture.

HELPFUL, BUT NOT REQUIRED:

* Experience with ML graph optimizations, parallel / distributed programming,

heterogeneous ML computation, and/or code generation.

* Exposure to MLIR, LLVM, and/or the Mojo programming language.

* Advanced degree in Computer Science or a related area is a plus.

WHAT MODULAR BRINGS TO THE TABLE:

* Amazing Team. We are a progressive and agile team with some of the industry’s

best engineering and product leaders.

* World-class Benefits. In order to attract the best, we need to offer the

best. Your benefits package may include comprehensive healthcare coverage,

retirement and savings programs, employee stock purchase opportunities, paid

time off, wellbeing resources, family support programs, and learning and

development opportunities. Please note that specific benefit packages may

vary based on your location, you can read more about benefits offered by

Qualcomm here [https://www.qualcomm.com/company/careers/benefits].

* Competitive Compensation. We offer very strong compensation packages,

including RSU grants. We want people to be focused on their best work and

believe in tailoring compensation plans to meet the needs of our workforce.

* Team Building Events. We organize regular team onsites and local meetups in

Los Altos, CA as well as different cities. Traveling 2-4 times a year is

expected for all roles.

Working at Modular will enable you to grow quickly as you work alongside

incredibly motivated and talented people who have high standards, possess a

growth mindset, and a purpose to truly change the world.

The estimated base salary range for this role to be performed in the US,

regardless of the state, is $216,000.00 - $324,000.00 USD.

The estimated base salary range for this role to be performed in Canada,

regardless of the province, is $207,000.00 - $310,400.00 CAD.

The salary for the successful applicant will depend on a variety of permissible,

non-discriminatory job-related factors, which include but are not limited to

education, training, work experience, business needs, or market demands. This

range may be modified in the future. The total compensation for a candidate will

also include annual target bonus, equity, and benefits, with equity making up a

significant portion of your total compensation.

For candidates who fall outside of the listed requirements, we nevertheless

encourage you to apply as we may have upcoming openings that are lower/higher

level than the ones advertised.

Which skills does this role require?

C++High-performance computingCPU optimizationGPU optimizationPerformance analysisProfiling toolsAI runtimeCompiler designKernel developmentDistributed programmingHeterogeneous computingML graph optimizationData localityBenchmarkingAPI developmentAI RuntimeCPUGPUMLIRLLVMMojoProfilingCompilerKernelsHeterogeneous computationAI platformSoftware engineeringOptimizationHardware accelerationMachine Learning

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.