Prepin
Log in
ECS Tech Inc

engineering opportunity

Cloud Platforms Engineer

The Cloud Platforms Engineer will design, deploy, and maintain highly available, fault-tolerant cloud infrastructure using Terraform and Infrastructure as Code principles. They will also implement monitoring, logging, and disaster recovery solutions while collaborating with cross-functional teams to ensure mission-critical workloads remain secure and scalable.

Fairfax, Virginia, United StateshybridFULL_TIME

Posted

About the role

What will you do at ECS Tech Inc?

Everforth ECS is seeking an experienced Cloud Platforms Engineer specializing in

reliability and resiliency to work in our Fairfax, VA office in a hybrid

capacity.

Everforth ECS is seeking an experienced Cloud Platforms Engineer specializing in

reliability and resiliency to design, operate, and continuously improve an Azure

Government cloud infrastructure supporting mission-critical workloads for

multiple coalition Mission Partner Network enclaves in support of the DoW

community. This position will have a strong focus on infrastructure reliability,

automation, High Availability (HA), Disaster Recovery (DR), and Infrastructure

as Code (IaC), with Terraform serving as a primary platform for provisioning and

managing cloud resources.

The Cloud Platforms Engineer will work across cloud infrastructure, platform

engineering, security, and application teams to ensure production environments

remain reliable, scalable, recoverable, secure, and operationally sustainable.

The ideal candidate combines a strong cloud architecture skill set with hands-on

operational experience and an automation-first approach to infrastructure

management.

Key Responsibilities

  • * Design, deploy, and maintain highly available, fault-tolerant cloud
  • infrastructure using Terraform and Infrastructure as Code principles.
  • * Architect and implement High Availability (HA) and Disaster Recovery (DR)
  • solutions, including infrastructure redundancy, automated failover, backup
  • and restoration, geographic resiliency, and recovery procedures.
  • * Manage the complete lifecycle of cloud infrastructure, including
  • provisioning, configuration, operating system and platform maintenance,
  • patching, upgrades, vulnerability remediation, and decommissioning.
  • * Develop automation for infrastructure deployment and routine operational
  • activities to reduce manual administration, configuration drift, and the
  • potential for human error.
  • * Implement and maintain monitoring, logging, alerting, and observability
  • capabilities to identify infrastructure degradation, capacity constraints,
  • performance issues, and potential service disruptions before they impact
  • users.
  • * Participate in incident response, troubleshooting, and root-cause analysis
  • for production infrastructure events and develop corrective actions to
  • prevent recurrence.
  • * Define and maintain infrastructure reliability standards, including Service
  • Level Indicators (SLIs), Service Level Objectives (SLOs), availability
  • targets, recovery time objectives (RTOs), and recovery point objectives
  • (RPOs).
  • * Develop, maintain, and regularly validate disaster recovery procedures
  • through recovery exercises, failover testing, and infrastructure restoration
  • testing.
  • * Evaluate cloud infrastructure capacity, performance, availability, and
  • scalability and recommend architectural or operational improvements.
  • * Partner with application development, cybersecurity, DevOps, and platform
  • engineering teams to establish standardized deployment patterns and resilient
  • cloud architectures.
  • * Maintain infrastructure documentation, operational procedures, architecture
  • diagrams, runbooks, and recovery procedures required to support production
  • environments.
  • * Provide technical leadership and guidance regarding cloud infrastructure
  • reliability, resiliency, automation, and operational best practices.
  • * Other duties, as assigned.
  • Please note, Salary is commensurate with skillset, qualifications, experience,
  • and educational background.

Qualifications

  • * U.S. Citizen.
  • * Active DoD Secret security clearance.
  • * Bachelor’s degree with 5+ years of related work experience.
  • * Ability to obtain a DoD 8140 IAT Level II Security+ (or higher) within 60
  • days of hire.
  • * Ability to work in a hybrid capacity in Fairfax, VA (up to 3 days in office).
  • * Ability to travel <20% throughout the lifespan of the Program to CONUS /
  • OCONUS customer sites and government installations.
  • * Strong experience designing, deploying, and supporting highly available,
  • fault-tolerant cloud infrastructure.
  • * Hands-on experience with Terraform and Infrastructure as Code (IaC)
  • principles for provisioning and managing cloud resources.
  • * Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture,
  • including redundancy, automated failover, backup and restoration, geographic
  • resiliency, and recovery planning.
  • * Experience managing the full cloud infrastructure lifecycle, including
  • provisioning, configuration, patching, upgrades, vulnerability remediation,
  • maintenance, and decommissioning.
  • * Strong infrastructure automation skills with an emphasis on reducing manual
  • administration, configuration drift, and operational error.
  • * Experience implementing and operating monitoring, logging, alerting, and
  • observability solutions for production infrastructure.
  • * Strong troubleshooting and diagnostic skills, including incident response,
  • root-cause analysis, and corrective action development.
  • * Understanding of Site Reliability Engineering concepts, including SLIs, SLOs,
  • availability targets, RTOs, and RPOs.
  • * Experience developing and validating disaster recovery procedures, including
  • failover exercises, recovery testing, and infrastructure restoration.
  • * Ability to evaluate infrastructure capacity, performance, scalability,
  • availability, and resiliency and recommend architectural or operational
  • improvements.
  • * Experience working collaboratively with application development,
  • cybersecurity, DevOps, and platform engineering teams.
  • * Ability to develop and maintain technical documentation, including
  • architecture diagrams, operational procedures, runbooks, and recovery
  • documentation.
  • * Strong understanding of cloud infrastructure security, vulnerability
  • management, and operational best practices.
  • * Demonstrated ability to provide technical leadership and guidance in
  • infrastructure reliability, resiliency, automation, and cloud operations.
  • * Experience supporting production, enterprise, regulated, or mission-critical
  • environments is highly desirable.
  • * Strong problem-solving and decision-making capabilities, with a proven
  • ability to weigh the relative costs and benefits of potential actions and
  • identify the most appropriate solution.
  • * Highly developed interpersonal and oral/written communication skills, with
  • the ability to effectively and professionally interact with a diverse set of
  • stakeholders (from peers to end-users to executive management).

Which skills does this role require?

Azure GovernmentTerraformInfrastructure as CodeHigh AvailabilityDisaster RecoveryCloud ArchitectureAutomationSite Reliability EngineeringMonitoringLoggingObservabilityIncident ResponseRoot-cause AnalysisVulnerability RemediationPlatform EngineeringSecurityCloud Platforms EngineerSLISLORTORPODoD Secret ClearanceSecurity+Mission Partner NetworkDevOpsCybersecurityCloud InfrastructureFault-tolerantVulnerability ManagementPatchingProvisioningConfiguration ManagementSystem AdministrationEnterprise EnvironmentsMission-criticalAzure

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.