Prepin
Log in
Instrumental Inc.

engineering opportunity

Site Reliability Engineer

You will operate, improve, and scale the company's AWS-based SaaS platform while focusing on reliability, automation, and observability. You will also participate in an on-call rotation to manage production incidents and engineer solutions to reduce operational complexity.

Palo Alto, California, United StatesremoteFULL_TIME

Posted

About the role

What will you do at Instrumental Inc.?

Instrumental builds the manufacturing acceleration platform behind the world’s

most complex electronics. We capture digital exhaust and engineering context

from assembly lines—images, test logs, BOM data, performance, repair cycles—and

our AI engines identify insights that are difficult or impossible for human

engineers to find. We accelerate the companies building the AI era by improving

manufacturing yield, throughput, and ramp. NVIDIA, Meta, Cisco, and their

manufacturing partners rely on Instrumental to accelerate new product

introduction and production.

The Instrumental platform collects, intelligently transforms, and contextually

presents manufacturing data to technical end-users, enabling them to optimize

their manufacturing process in real-time. Our core technology is proprietary ML

algorithms, packaged in an accessible, user-centric user interface—we believe we

must have both the best technology and the best access to that technology to

win.

As a Site Reliability Engineer, you’ll operate, improve, and scale our AWS-based

SaaS platform. You’ll combine hands-on production operations with engineering,

focusing on reliability, automation, observability, and operational excellence.

You’ll participate in a bi-weekly on-call rotation, but the goal isn’t simply to

keep systems running—it’s to continuously engineer away the operational

complexity that comes with scaling our platform and customer base.

Requirements

  • * 3–4 years of experience in Site Reliability Engineering, DevOps, Cloud
  • Operations, Platform Engineering, or Systems Engineering supporting
  • production SaaS environments.
  • * Strong hands-on experience with AWS, including EC2, VPC, IAM, RDS, ECS, and
  • S3.
  • * Experience managing infrastructure using Terraform or other Infrastructure as
  • Code technologies.
  • * Experience designing and supporting CI/CD pipelines using GitHub Actions,
  • Jenkins, GitLab CI/CD, or similar platforms.
  • * Strong experience with monitoring and observability tools, preferably
  • Datadog, including dashboards, alerting, logging, and APM.
  • * Experience with Docker and Kubernetes.
  • * Scripting experience with Python and/or Bash.
  • * Experience supporting production environments through an on-call rotation,
  • including incident response and root cause analysis.
  • * Proven ability to take ownership of production issues and drive them through
  • investigation, remediation, and long-term resolution.

Who You Are

  • * Dead serious about performance, scalability, and reliability (PSR): You care
  • deeply about how systems behave in the real world and continuously look for
  • ways to make them more reliable, scalable, observable, and supportable.
  • * Automation, automation, automation: If something is repetitive, manual, or
  • error-prone, your first instinct is to automate it and make it disappear.
  • * An engineer at heart: You don’t want to repeatedly fight the same fires. You
  • look for the underlying cause and build durable engineering solutions that
  • reduce operational toil and technical debt.
  • * Strong systems thinker: You understand how infrastructure, applications,
  • networks, deployments, monitoring, and people interact—and can troubleshoot
  • complex production issues across those boundaries.
  • * Collaborative and reliable: You partner closely with software engineers to
  • make services production-ready, improve operational workflows, and build
  • reliability into systems before they become problems.
  • * Comfortable with growth and ambiguity: You’re comfortable making good
  • decisions without perfect information and adapting as the platform, customer
  • base, and company scale quickly.

Nice to Have

  • * Experience working in a high-growth B2B SaaS environment.
  • * Experience implementing SRE practices such as SLIs, SLOs, and error budgets.
  • * Experience building internal tooling and automation to eliminate operational
  • toil.
  • * Experience supporting multi-region AWS environments.
  • * AWS cost optimization or FinOps experience.
  • * Network, application security, and compliance experience.
  • * Experience introducing AI tools or processes into engineering and operational
  • workflows.
  • This position requires access to items and data that are developed under U.S.
  • government contracts and subject to dissemination controls that limit access to
  • U.S. citizens only.
  • We’re a growing team that works collaboratively, is supportive of each other,
  • and is highly energized by the opportunity for a large impact. We actively work
  • to promote an inclusive environment, valuing passion and the ability to learn.
  • You’re encouraged to apply even if your experience doesn’t precisely match the
  • job description!
  • The following is a representative annual base salary range for this position
  • within the Bay Area: $140,000-$165,000. Job level and salary opportunities are
  • evaluated through our interview process – we review the experience, knowledge,
  • skills, and abilities of each applicant.
  • Instrumental is proud to offer a highly-rated variety of benefits, including
  • health, vision, dental, commuter plans, and parental leave.
  • At Instrumental, protecting company and customer information is a shared
  • responsibility. Employees are expected to comply with company engineering,
  • security, access control, and privacy policies, and promptly report suspected
  • security incidents or policy violations.

Which skills does this role require?

Site Reliability EngineeringTerraformInfrastructure as CodeCI/CDDockerKubernetesPythonBashDatadogObservabilityIncident responseSystems engineeringCloud operationsDevOpsGitHub ActionsJenkinsGitLab CI/CDProduction operationsRoot cause analysisScalabilityReliabilityManufacturingMachine learningCloud computingMachine Learning

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.