About the role
What will you do at Modular, a Qualcomm company?
ABOUT THE ROLE:
ML developers today face significant friction when deploying trained models.
They work in a fragmented space with incomplete, patchwork solutions that
require extensive performance tuning and model-specific optimizations. At
Modular, we are building the next-generation AI platform that will radically
improve how developers build and deploy AI models.
A core part of this offering is a platform that enables customers to achieve
state-of-the-art performance across model families and frameworks. As an AI
Runtime Engineer, you will own a runtime that operates on various CPU and GPU
hardware platforms, optimizing performance for diverse customer AI models.
LOCATION: Candidates based in the US or Canada are welcome to apply. You can
work in our office in Los Altos, CA or remotely from home. Onboarding for new
hires is conducted in-person in our Los Altos, CA office.
WHAT YOU WILL DO:
* Design and develop runtime and cross-stack optimizations to improve CPU and
GPU efficiency, addressing issues such as CPU overhead, caching, and data
locality across multiple devices.
* Port the Modular runtime stack to new hardware platforms and develop an API
to streamline this process.
* Collaborate with the compiler, kernels, serving, and models teams to design
core technologies that achieve state-of-the-art end-to-end performance on
various CPU and GPU hardware.
* Collaborate with the customer success team and engage with customers to
understand their performance requirements and use cases.
* Collaborate with tooling and infrastructure teams to design systems for
automated performance analysis and benchmarking.
WHAT YOU BRING TO THE TABLE:
* 5+ years of experience working on high-performance computing systems.
* Experience in C++ programming and complex software systems.
* Experience with CPU or GPU runtime optimizations and performance analysis on
CPUs, GPUs, or AI accelerators.
* Proficiency with one or more profiling tools (CPU or GPU).
* Creativity and curiosity for solving complex problems, a team-oriented
attitude that enables you to work well with others, and alignment with our
culture.
HELPFUL, BUT NOT REQUIRED:
* Experience with ML graph optimizations, parallel / distributed programming,
heterogeneous ML computation, and/or code generation.
* Exposure to MLIR, LLVM, and/or the Mojo programming language.
* Advanced degree in Computer Science or a related area is a plus.
WHAT MODULAR BRINGS TO THE TABLE:
* Amazing Team. We are a progressive and agile team with some of the industry’s
best engineering and product leaders.
* World-class Benefits. In order to attract the best, we need to offer the
best. Your benefits package may include comprehensive healthcare coverage,
retirement and savings programs, employee stock purchase opportunities, paid
time off, wellbeing resources, family support programs, and learning and
development opportunities. Please note that specific benefit packages may
vary based on your location, you can read more about benefits offered by
Qualcomm here [https://www.qualcomm.com/company/careers/benefits].
* Competitive Compensation. We offer very strong compensation packages,
including RSU grants. We want people to be focused on their best work and
believe in tailoring compensation plans to meet the needs of our workforce.
* Team Building Events. We organize regular team onsites and local meetups in
Los Altos, CA as well as different cities. Traveling 2-4 times a year is
expected for all roles.
Working at Modular will enable you to grow quickly as you work alongside
incredibly motivated and talented people who have high standards, possess a
growth mindset, and a purpose to truly change the world.
The estimated base salary range for this role to be performed in the US,
regardless of the state, is $216,000.00 - $324,000.00 USD.
The estimated base salary range for this role to be performed in Canada,
regardless of the province, is $207,000.00 - $310,400.00 CAD.
The salary for the successful applicant will depend on a variety of permissible,
non-discriminatory job-related factors, which include but are not limited to
education, training, work experience, business needs, or market demands. This
range may be modified in the future. The total compensation for a candidate will
also include annual target bonus, equity, and benefits, with equity making up a
significant portion of your total compensation.
For candidates who fall outside of the listed requirements, we nevertheless
encourage you to apply as we may have upcoming openings that are lower/higher
level than the ones advertised.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Research Engineer, Responsible Frontier AI Research, DeepMind
Google · New York, New York, United States
AI Outcome Customer Engineer, Forward Deployed Engineering
Google · Atlanta, Georgia, United States
Senior SRAM Circuit Design Engineer - AI & HPC (7708)
TSMC · San Jose, California, United States
AI Risk Engineer
Bright Vision Technologies · Columbus, Ohio, United States
Applied AI Design Engineer (100 % remote) (m/f/d)
EWOR · Capon Bridge, West Virginia, United States
VP – Distinguished Engineer of Generative AI Engineering
Slate Auto · United States
Role information can change. Confirm current details on the original application page.
