About the role
What will you do at Hippocractic AI?
About the RoleWe're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models.We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge.
You'll build the GPU management and scheduling platform that sits at the center of it — collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling.
The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts.This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.What You'll DoDesign and build our GPU management and scheduling platform — the system that decides when, where, and how inference calls run across a fleet of ~30 models on heterogeneous hardwareBuild the metrics pipeline that collects GPU load and utilization data, and the logic that turns those signals into decisionsImplement admission control to protect capacity — deciding when to accept, queue, or shed inference requests so we operate within fleet limitsBuild autoscaling that adjusts the number of model replicas in response to real-time demand and utilizationDevelop cloud orchestration systems and operators in Python and Go to manage the model fleetArchitect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or AzureDesign and build infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class softwareStand up and maintain monitoring, logging, and alerting that keep the platform reliable and performantDevelop and enforce security and compliance policies appropriate to a healthcare AI platformPartner with engineers and research scientists to diagnose and resolve complex infrastructure, deployment, and operational issuesMentor engineers and raise the technical bar across the teamWhat You BringMust-Have10+ years of professional experience across site reliability / DevOps engineering and software engineeringComputer Science Degree Required from a top CS program.Strong software engineering fundamentals — you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf toolsExperience designing systems that make decisions from operational metrics — collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission controlDeep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)Hands-on production experience with at least one major cloud platform (AWS, GCP, or Azure)Strong knowledge of containerization and orchestration (Docker, Kubernetes)Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)Excellent problem-solving skills and the ability to work both independently and collaborativelyStrong communication and interpersonal skillsNice-to-HaveExperience managing GPU fleets or scheduling workloads across heterogeneous acceleratorsFamiliarity with ML inference serving and model deployment (e.g.
Triton, KServe, Ray Serve, or similar)Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)Experience implementing HIPAA and SOC 2 complianceExperience operating in an HPC environmentBachelor's or Master's in Computer Science, Computer Engineering, or a related fieldJoin our team at Hippocratic AI and help shape the future of clinically safe, production-grade AI systems.Why Join Hippocratic AIReinvent healthcare with AI that puts safety first.
We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St.
Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others.Build alongside the best in healthcare and AI.
Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative.Equal OpportunityHippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law.
We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact [email protected] be aware of recruitment scams impersonating Hippocratic AI.
All recruiting communication will come from @hippocraticai.com email addresses. We will never request payment or sensitive personal information during the hiring process.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior Software Engineer, Infrastructure
HUD · San Francisco / Singapore
Senior Site Reliability Engineer
Replit · United States
Senior Software Engineer - Infrastructure Security
Emergent · Bangalore
Senior Enterprise Infrastructure Engineer
Boku · London, GB
Technical Account Manager (OEM)
Aiven · Austin, TX, US
Tech Lead - Code Plane [IC5]
Sourcegraph · Remote
Role information can change. Prepin can help you prepare, but does not submit an application for this role.