Prepin
Log in
Databricks

engineering opportunity

Senior GenAI Research Engineer - Optimization and Kernels

As a research engineer on the Scaling Team, you will drive performance improvements through advanced optimization techniques and design high-performance GPU kernels for training workloads. You will also implement distributed training frameworks for large language models and optimize end-to-end training workflows.

Mountain View, California, United StatesremoteFULL_TIME

Posted

About the role

What will you do at Databricks?

At Databricks, we are obsessed with enabling data teams to solve the world’s

toughest problems, from security threat detection to cancer drug development. We

do this by building and running the world’s best data and AI platform so our

customers can focus on the high-value challenges that are central to their own

missions.

The Databricks AI Research organization enables companies to develop AI models

and agents using their own data, with technologies ranging from post-training

open source LLMs to developing advanced multi-agent architectures. Databricks AI

is committed to the belief that a company’s AI models and agents are just as

valuable as any other core IP, and that high-quality AI should be available to

all.

Job Description

As a Sr. Research Engineer on the Scaling team, you will be responsible for

keeping up with the latest developments in deep learning and advancing the

scientific frontier by creating new techniques that go beyond the state of the

art. You will work together on a collaborative team of researchers and engineers

with diverse backgrounds and technical training. And most importantly, you will

love our customers: our goal is to make our customers successful in applying

state-of-the-art LLMs and AI systems, and we encode our scientific expertise

into our products to make that possible.

The Impact you will have

As a research engineer on the Scaling Team at Databricks, you will:

* Drive performance improvements through advanced optimization techniques

including kernel fusion, mixed precision, memory layout optimization, tiling

strategies, and tensorization for training-specific patterns

* Design, implement, and optimize high-performance GPU kernels for training

workloads (e.g., attention mechanisms, custom layers, gradient computation,

activation functions) targeting NVIDIA architectures

* Design and implement distributed training frameworks for large language

models, including parallelism strategies (data, tensor, pipeline, ZeRO-based)

and optimized communication patterns for gradient synchronization and

collective operations

* Profile, debug, and optimize end-to-end training workflows to identify and

resolve performance bottlenecks, applying memory optimization techniques like

activation checkpointing, gradient sharding, and mixed precision training.

What We Look for

* BS/MS/PhD in Computer Science or related field with hands-on experience

writing and tuning CUDA kernels for ML training applications, or hands-on

experience in distributed training frameworks (PyTorch DDP, DeepSpeed,

Megatron-LM, FSDP)

* Strong understanding of NVIDIA GPU architecture (memory hierarchy, tensor

cores, warp scheduling, SM occupancy) and proficiency with CUDA

debugging/profiling tools (Nsight, NVProf)

* Deep understanding of parallelism techniques and memory optimization

strategies for large-scale model training, with proven ability to debug and

optimize distributed workloads

* Strong software engineering skills in Python and PyTorch, with experience

supporting production training workflows and knowledge of LLM training

dynamics including hyperparameter tuning and optimization strategies.

Pay Range Transparency

Databricks is committed to fair and equitable compensation practices. The pay

range(s) for this role is listed below and represents the expected salary range

for non-commissionable roles or on-target earnings for commissionable roles.

Actual compensation packages are based on several factors that are unique to

each candidate, including but not limited to job-related skills, depth of

experience, relevant certifications and training, and specific work location.

Based on the factors above, Databricks anticipates utilizing the full width of

the range. The total compensation package for this position may also include

eligibility for annual performance bonus, equity, and the benefits listed above.

For more information regarding which range your location is in visit our page

here

[https://www.databricks.com/sites/default/files/2024-08/us-pay-zone-mapping.pdf].

Local Pay Range

$166,000—$225,000 USD

About Databricks

Databricks is the data and AI company. More than 10,000 organizations worldwide

— including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 —

rely on the Databricks Data Intelligence Platform to unify and democratize data,

analytics and AI. Databricks is headquartered in San Francisco, with offices

around the globe and was founded by the original creators of Lakehouse, Apache

Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter

[https://twitter.com/databricks], LinkedIn

[https://www.linkedin.com/company/databricks] and Facebook

[https://www.facebook.com/databricksinc].

Benefits

At Databricks, we strive to provide comprehensive benefits and perks that meet

the needs of all of our employees. For specific details on the benefits offered

in your region click here

[https://docs.google.com/document/d/154un3e8Xav4BceOSlcYFZRGEuQI54xMxVydRwQn54eQ/edit? usp=sharing].

Our Commitment to Diversity and Inclusion

At Databricks, we are committed to fostering a diverse and inclusive culture

where everyone can excel. We take great care to ensure that our hiring practices

are inclusive and meet equal employment opportunity standards. Individuals

looking for employment at Databricks are considered without regard to age,

color, disability, ethnicity, family or marital status, gender identity or

expression, language, national origin, physical and mental ability, political

affiliation, race, religion, sexual orientation, socio-economic status, veteran

status, and other protected characteristics.

Compliance

If access to export-controlled technology or source code is required for

performance of job duties, it is within Employer's discretion whether to apply

for a U.S. government license for such positions, and Employer may decline to

proceed with an applicant on this basis alone.

Which skills does this role require?

CUDADeep LearningGPU KernelsOptimization TechniquesDistributed TrainingNVIDIA ArchitecturePythonPyTorchMemory OptimizationParallelismGradient ComputationActivation FunctionsMixed PrecisionDebuggingProfilingPerformance BottlenecksLarge Language ModelsTraining WorkloadsAttention MechanismsCustom LayersHyperparameter TuningTensorizationKernel FusionMemory Layout OptimizationTiling StrategiesZeRO-basedMachine LearningLLMs

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.