Prepin
Log in
Mistral

engineering opportunity

NEW - Systems Engineer, HPC

Design, operate, and scale the infrastructure powering AI platforms, focusing on large-scale Linux environments and HPC clusters. Automate operational tasks and collaborate with research teams to optimize workload management and system reliability.

Palo Alto, California, United StatesremoteFULL_TIME

Posted

About the role

What will you do at Mistral?

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East.

We are creative, low-ego and team-spirited.

The Role

As a Systems Engineer / System Administrator in the Compute team, you will help design, operate, and scale the infrastructure powering Mistral’s AI platforms. This role exists to ensure our large-scale Linux environments, HPC clusters, and cloud systems run reliably, enabling cutting-edge research and production workloads. You’ll collaborate closely with infrastructure, HPC, and research teams to bridge the gap between users and systems, solving complex problems in a rapidly scaling environment.

Your work will directly impact the performance, reliability, and scalability of systems handling petabyte-scale data and thousands of nodes.

What You Will Do

  • Operate and maintain large-scale Linux environments, including bare metal, clusters, and cloud infrastructure.
  • Monitor system health, troubleshoot incidents, and ensure high availability for production and research workloads.
  • Scale clusters toward hundreds to thousands of nodes while improving performance and resource utilization.
  • Automate operational tasks using tools like Python, Bash, Ansible, or Terraform to streamline deployment and system lifecycle management.
  • Contribute to system design and architecture decisions to enhance reliability and scalability.
  • Work with job schedulers (e.g., Slurm) to optimize workload management.
  • Collaborate with HPC, infrastructure, and research teams to align systems with user needs.
  • Act as a bridge between technical teams and end-users to resolve cross-functional challenges.

What We're Looking For

  • Strong Linux systems administration experience in large-scale environments (HPC clusters or cloud infrastructure).
  • Experience with job schedulers (e.g., Slurm) and troubleshooting across systems, hardware, and networks.
  • Familiarity with containers/orchestration (e.g., Kubernetes), storage systems (e.g., Ceph, Lustre, NFS), or networking fundamentals (Ethernet, InfiniBand).
  • Proficiency with Infrastructure as Code or automation tooling (e.g., Ansible, Terraform).
  • Exposure to GPU or AI/ML workloads is a plus.
  • Pragmatic problem-solving skills with the ability to operate in fast-scaling environments.
  • Comfortable working across multiple domains with a "Swiss army knife" mindset.
  • Low-ego, collaborative, and hands-on approach to work.

What we offer

  • We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance.

Benefits

vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Which skills does this role require?

Linux Systems AdministrationHPC ClustersSlurmKubernetesCephLustreNFSInfiniBandAnsibleTerraformPythonBashGPU WorkloadsInfrastructure as CodeCloud InfrastructureSystem ArchitectureLinuxHPCEthernetGPUAI/MLBare MetalSystem AdministrationScalabilityHigh AvailabilityResource UtilizationWorkload ManagementPetabyte-scale DataMachine Learning

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.