Prepin
Log in
Varo Bank

engineering opportunity

Site Reliability Engineer (Contractor)

The Site Reliability Engineer will manage and scale EKS clusters, maintain infrastructure as code, and support data platforms like Kafka. They will also participate in incident response, on-call rotations, and automate operational tasks to improve system observability.

Salt Lake City, Utah, United StatesremoteFULL_TIME

Posted

About the role

What will you do at Varo Bank?

Varo is an entirely new kind of bank. All digital, mission-driven, FDIC insured

and designed for the way our customers live their lives. A bank for all of us.

ABOUT THE ROLE

Varo’s SRE team is well established, designing, building, and running

large-scale, distributed, fault-tolerant systems that power most of Varo's

operations. We live and breathe AWS and Kubernetes, having an open source first

and result oriented mindset.

We are an automation and observability focused team and we strive to automate

ourselves out of manual / remedial tasks. We monitor and create dashboards to

promote a data-driven approach to scale out our platform.

On a typical day, members of our team are hands-on scaling-out production

infrastructure, building out CI/CD pipelines, and brainstorming with developers

on how to make things better. We collectively strive to build and maintain a

rapid-feedback platform that enables our engineers to accomplish their own goals

instead of creating friction.

RESPONSIBILITIES

* EKS & Karpenter Management: Manage, upgrade, and autoscale EKS clusters

across multiple environments (SIT, UAT, Prod) and AWS accounts.

* Infrastructure as Code & GitOps: Write Terraform modules and Helm charts to

support GitOps workflows using ArgoCD and GitLab CI/CD pipelines.

* Kafka & Data Platform Support: Maintain and troubleshoot Kafka (MSK)

clusters, including broker health, connectors, and CDC pipelines.

* Observability & Cost Control: Improve observability using Prometheus, Thanos,

Grafana, and ELK while proactively identifying cloud cost-optimization

opportunities.

* Automation & AIOps: Automate operational tasks with Python and leverage AI/ML

techniques for predictive alerting and intelligent runbooks.

* Service Desk & Support: Handle Platform Service Desk requests, including

Terraform merge request reviews, access management, and deployment support.

* Incident Response & On-Call: Participate in the production on-call rotation,

support incident response, and contribute to blameless post-mortems.

SKILLS & QUALIFICATION

* Experience: 3+ years of experience in an SRE, DevOps, or Infrastructure

Engineering role, with the ability to work independently and manage multiple

workstreams.

* AWS Expertise: Strong hands-on experience with core AWS services, including

EKS, EC2, RDS Aurora, MSK, S3, IAM, VPC, and Direct Connect.

* Kubernetes & Deployment: Deep production experience with Kubernetes

(upgrades, networking, RBAC) alongside Helm and GitOps tools like ArgoCD.

* Infrastructure as Code: Advanced proficiency with Terraform, including

writing modules and managing multi-account/multi-environment states.

* Data Infrastructure: Experience supporting and maintaining data platforms

such as Airflow, Databricks, EMR, Kafka/MSK, or CDC pipelines.

* Networking & Automation: Solid understanding of networking (VPCs, security

groups, Istio, DNS) paired with strong Python scripting skills for tooling

and automation.

* Observability & AI: Experience managing observability stacks (Prometheus,

Grafana, ELK) and effectively leveraging AI/LLM tools for automation and

incident analysis.

*

Nice to Have

  • Experience with Karpenter and KEDA, GitLab CI/CD pipeline
  • experience, Hashicorp Vault for secrets management.
  • For cash compensation, we set standard ranges for all US-based roles based on
  • function, level, and geographic location, benchmarked against similar-stage
  • growth companies. Final offer amounts are determined by multiple factors as well
  • as candidate experience and expertise and may vary from the identified range.
  • Varo is an equal opportunity employer. Varo embraces diversity and we are
  • committed to building teams that represent a variety of backgrounds,
  • perspectives, and skills. All applicants will be considered for employment
  • without attention to race, color, religion, sex, sexual orientation, gender
  • identity, national origin, veteran or disability status.
  • Beware of fraudulent job postings!
  • Varo will never ask for payment to process documents, refer you to a third party
  • to process applications or visas, or ask you to pay costs. Never send money to
  • anyone suggesting they can provide work with Varo. If you suspect you have
  • received a phony offer, please e-mail [email protected] with the pertinent
  • information and contact information.
  • Notice at Collection for Employees and Applicants:
  • Varo Candidate Privacy Notice
  • [https://assets.ctfassets.net/x6cbfr3jz6wz/6GyZ5C54N0CbJ7Y5CMt26H/0b5aeaffc9bbcaa0b53e524dcd18a817/Job_Applicant_Privacy_Notice_April_2026.pdf]

Which skills does this role require?

AWSKubernetesTerraformPythonEKSKafkaGitOpsArgoCDPrometheusGrafanaELKInfrastructure as CodeNetworkingObservabilityIncident ResponseSite Reliability EngineerHelmMSKThanosAIOpsVPCIstioDNSAirflowDatabricksEMRCDCCloud Cost-OptimizationC#Machine LearningLLMs

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.