Prepin
Log in
ECS Tech Inc

engineering opportunity

Senior Site Reliability Engineer

The Senior SRE will define and implement SRE practices to ensure the reliability, availability, and performance of critical production environments. They will also manage logging, monitoring, and alerting solutions while collaborating with cross-functional teams to integrate observability into the software development lifecycle.

Fairfax, Virginia, United StatesremoteFULL_TIME

Posted

About the role

What will you do at ECS Tech Inc?

Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely.

Everforth ECS is seeking talented professionals to join our successful and

growing team in building the next-generation Continuous Diagnostics and

Mitigation (CDM) Cyber data solution. The CDM Program is the Cybersecurity and

Infrastructure Security Agency’s (CISA) dynamic approach to strengthening the

cybersecurity of Federal networks and systems through better awareness and

visibility into their security posture and cyber threats. ECS is responsible for

designing, building, deploying, operating, and maintaining a complete ‘Data

Services’ solution which includes the collection, normalization, visualization,

and sharing of cyber data from more than 100 Federal agencies. The CDM Data

Services product is an integrated suite of multiple Commercial Off the Shelf

(COTS) products, software configuration packages, and custom code which work

together to operate as an integrated solution tailored to meet Department of

Homeland Security (DHS) requirements. 

We are seeking professionals who thrive in a dynamic, fast-paced, and highly

collaborative environment where problem-solving, critical thinking, and a

holistic approach to serving the mission are key. Our program operates within

the Scaled Agile Framework (SAFe). An aptitude and enthusiasm for continuous

learning, improvement, and cyber security is a must!

Role &

Responsibilities

  • ECS is seeking a talented Senior Site Reliability Engineer (SRE) to play a key
  • role in defining, implementing, and growing our SRE practice to ensure the
  • reliability, availability, and performance of our critical production
  • environments.
  • The Senior SRE will contribute to a culture of continuous
  • improvement, identifying areas for enhancement, and driving initiatives to
  • improve system reliability, scalability, and efficiency.
  • The successful candidate will have demonstrated hands-on experience designing,
  • implementing, and maintaining solutions to ensure that systems,
  • including infrastructure and applications, are resilient, highly available, and
  • performant. The Senior SRE will also play a critical role in defining and
  • measuring the Service Level Objectives (SLOs) and Service Level Indicators
  • (SLIs) for our solution.
  • The Senior SRE will be responsible for setting up
  • comprehensive logging, monitoring, and alerting solutions using the
  • Elastic stack and other tools as necessary to ensure the continuous performance
  • of services. Additionally, they will respond to incidents, perform root cause
  • analyses, and implement solutions to prevent reoccurrences. The Senior SRE will
  • work in close collaboration with other SRE team
  • members, developers, testers, infrastructure engineers, DevOps engineers, and
  • other stakeholders to integrate reliability and observability into the software
  • development lifecycle.
  • Salary Range: $118,000 - $177,000
  • General Description of Benefits [https://ecstech.com/careers/benefits]

Qualifications

  • * Must be a US citizen with the ability to obtain Public Trust Suitability.
  • * 6+ years of experience as a Site Reliability Engineer (SRE) or equivalent
  • * 6+ years of demonstrated experience designing, implementing,
  • and maintaining observability solutions to include logging, monitoring, and
  • alerting
  • * 6+ years of hands-on experience with SRE tools (e.g., Elastic, Prometheus,
  • Grafana, Splunk, etc.)
  • * 3+ years defining and measuring SLOs and SLIs
  • * 3+ years of relevant experience using cloud platforms (AWS GovCloud
  • preferred)
  • * 3+ years of hands-on programming or scripting (e.g., Python, Bash, etc.)
  • * Strong knowledge of microservices, containerization, and orchestration tools
  • (Docker, Kubernetes)
  • * Proven ability to collaborate with cross-functional teams (development,
  • testing, and product) to integrate reliability and observability into the
  • software development lifecycle
  • * Strong problem-solving and analytical skills
  • * Proactive, detail-oriented approach to identifying inefficiencies and
  • implementing improvements.
  • * Proficient in developing Synthetic monitoring scripts using typescript.

Which skills does this role require?

Site Reliability EngineeringElastic StackPrometheusGrafanaSplunkAWS GovCloudService Level ObjectivesService Level IndicatorsCybersecuritySite Reliability EngineerContinuous Diagnostics and MitigationInfrastructureSLOSLIScaled Agile FrameworkSAFeCloud ComputingLoggingAlertingRoot Cause AnalysisDevOpsFederal SystemsPublic TrustData ServicesAWSAgile

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.