About the role
What will you do at ECS Tech Inc?
Everforth ECS is seeking a Cloud Site Reliability Engineer (SRE) to work in our
Arlington, VA office/remotely.
Our Philosophy
We believe the job of an SRE is to engineer the cloud to run itself. That means
writing software and automation that lets systems detect and recover from
failure on their own, rather than relying on someone to notice an alert and
manually fix it. When something breaks, self-healing comes first, deep
root-cause debugging happens after service is restored, not instead of it. We’re
looking for someone who automates the operational task by default, not documents
the runbook for doing it by hand.
About the Role
This role owns reliability and operational readiness for production systems
across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define
what “reliable enough” looks like for our services, build the automation that
gets us there, and do it all on an infrastructure-as-code (IaC) foundation.
Responsibilities
- Self-Healing Operations
- * Design and build automated remediation so systems detect, respond to, and
- recover from failure without manual intervention
- * Shift the team’s posture from “is it running, how do we fix it” to “how do we
- make it fix itself”
- * Automate service restoration first; investigate root cause after
- Uptime Goals & Reliability
- * Define reasonable, data-driven SLOs and error budgets for critical services
- alongside the teams that own them
- * Use live metrics to decide what’s “reliable enough” and where to invest next
- Infrastructure
- * Enforce infrastructure-as-code and configuration-as-code, no manual tech
- change
- * Own Terraform standards and reusable modules adopted across programs
- * Drive a containerization-first approach with production-scale Kubernetes
- (multi-tenancy, security policies, advanced scheduling)
- * Set CI/CD and pipeline-as-code standards, including progressive delivery
- Observability & Incidents
- * Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives
- automation the signal it needs to self-correct
- * Own the incident framework: escalation, restoration, root cause analysis, and
- post-incident review that closes the loop with more automation
- Collaboration & Leadership
- * Partner with development and contractor teams leads to embed reliability and
- automation across the software
- * Mentor engineers toward this same automation-first philosophy
- * Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure
- and automation
- Salary Range: $130,000 - $180,000
- General Description of Benefit [https://ecstech.com/careers/benefits]
Qualifications
- * Bachelor’s degree in Computer Science, Information Technology, or related
- field (or equivalent practical experience)
- * 5+ years of SRE experience (or equivalent), with demonstrated technical
- leadership
- * 10 years of general work experience
- * Track record building self-healing/auto-remediating systems, not just
- dashboards
- * Jenkins experience
- * Expert AWS knowledge, GovCloud experience strongly preferred
- * Deep Kubernetes and Terraform expertise at production scale
- * Strong software engineering background (Python and/or Go)
- * Experience operating observability platforms (Grafana, Splunk, Prometheus,
- Loki, etc.)
- * Proven incident command and postmortem experience
- * Strong communication skills across technical and federal leadership
- audiences
- * Ability to obtain/maintain required government clearance or suitability
- (CAC/PIV as applicable)
- * US Citizenship
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior Infrastructure Engineer - Data Protection
USAA · Tampa, Florida, United States
AI Infrastructure AI/ML Engineer (100 % remote) (m/f/d)
EWOR · Capon Bridge, West Virginia, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Infrastructure Engineer I - Data Protection
USAA · Tampa, Florida, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Neural Data Infrastructure Engineer
Blackrock Neurotech · Salt Lake City, Utah, United States
Role information can change. Confirm current details on the original application page.
