About the role
What will you do at ECS Tech Inc?
Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely.
Everforth ECS is seeking talented professionals to join our successful and
growing team in building the next-generation Continuous Diagnostics and
Mitigation (CDM) Cyber data solution. The CDM Program is the Cybersecurity and
Infrastructure Security Agency’s (CISA) dynamic approach to strengthening the
cybersecurity of Federal networks and systems through better awareness and
visibility into their security posture and cyber threats. ECS is responsible for
designing, building, deploying, operating, and maintaining a complete ‘Data
Services’ solution which includes the collection, normalization, visualization,
and sharing of cyber data from more than 100 Federal agencies. The CDM Data
Services product is an integrated suite of multiple Commercial Off the Shelf
(COTS) products, software configuration packages, and custom code which work
together to operate as an integrated solution tailored to meet Department of
Homeland Security (DHS) requirements.â¯
We are seeking professionals who thrive in a dynamic, fast-paced, and highly
collaborative environment where problem-solving, critical thinking, and a
holistic approach to serving the mission are key. Our program operates within
the Scaled Agile Framework (SAFe). An aptitude and enthusiasm for continuous
learning, improvement, and cyber security is a must!
Role &
Responsibilities
- ECS is seeking a talented Senior Site Reliability Engineer (SRE) to play a key
- role in defining, implementing, and growing our SRE practice to ensure the
- reliability, availability, and performance of our critical production
- environments.
- The Senior SRE will contribute to a culture of continuous
- improvement, identifying areas for enhancement, and driving initiatives to
- improve system reliability, scalability, and efficiency.
- The successful candidate will have demonstrated hands-on experience designing,
- implementing, and maintaining solutions to ensure that systems,
- including infrastructure and applications, are resilient, highly available, and
- performant. The Senior SRE will also play a critical role in defining and
- measuring the Service Level Objectives (SLOs) and Service Level Indicators
- (SLIs) for our solution.
- The Senior SRE will be responsible for setting up
- comprehensive logging, monitoring, and alerting solutions using the
- Elastic stack and other tools as necessary to ensure the continuous performance
- of services. Additionally, they will respond to incidents, perform root cause
- analyses, and implement solutions to prevent reoccurrences. The Senior SRE will
- work in close collaboration with other SRE team
- members, developers, testers, infrastructure engineers, DevOps engineers, and
- other stakeholders to integrate reliability and observability into the software
- development lifecycle.
- Salary Range: $118,000 - $177,000
- General Description of Benefits [https://ecstech.com/careers/benefits]
Qualifications
- * Must be a US citizen with the ability to obtain Public Trust Suitability.
- * 6+ years of experience as a Site Reliability Engineer (SRE) or equivalent
- * 6+ years of demonstrated experience designing, implementing,
- and maintaining observability solutions to include logging, monitoring, and
- alerting
- * 6+ years of hands-on experience with SRE tools (e.g., Elastic, Prometheus,
- Grafana, Splunk, etc.)
- * 3+ years defining and measuring SLOs and SLIs
- * 3+ years of relevant experience using cloud platforms (AWS GovCloud
- preferred)
- * 3+ years of hands-on programming or scripting (e.g., Python, Bash, etc.)
- * Strong knowledge of microservices, containerization, and orchestration tools
- (Docker, Kubernetes)
- * Proven ability to collaborate with cross-functional teams (development,
- testing, and product) to integrate reliability and observability into the
- software development lifecycle
- * Strong problem-solving and analytical skills
- * Proactive, detail-oriented approach to identifying inefficiencies and
- implementing improvements.
- * Proficient in developing Synthetic monitoring scripts using typescript.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior Infrastructure Engineer - Data Protection
USAA · Tampa, Florida, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
AI Infrastructure AI/ML Engineer (100 % remote) (m/f/d)
EWOR · Capon Bridge, West Virginia, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Neural Data Infrastructure Engineer
Blackrock Neurotech · Salt Lake City, Utah, United States
Role information can change. Confirm current details on the original application page.
