About the role
What will you do at ECS Tech Inc?
Everforth ECS is seeking a Senior ML Observability Engineer to work in the
National Capital Region covering the Pentagon, Falls Church, and Fairfax. Please
Note: This position is contingent upon contract award.
The War Data Platform (WDP) is a key initiative within the U.S. Department of
War's (DoW) AI-First strategy introduced in early 2026. The WDP focuses on
operational warfighting data and aims to accelerate the deployment of artificial
intelligence (AI) on the battlefield. The WDP extends to Unclassified, Secret,
and Top Secret environments, and supports collaboration between Combatant
Commands, Joint Staff directorates, Senior Executive Service leaders, and
operational analysts.
The Senior ML Observability Engineer architects and governs the instrumentation
and telemetry infrastructure needed to ensure production AI and machine learning
models deployed across WDP's multi-enclave environment perform reliably and
securely at mission scale. This role is essential to maintaining real-time
visibility into model behavior, pipeline execution, and cross-domain access
interactions in direct support of Combatant Command and Joint Staff
decision-making needs.
• Designs, implements, and governs observability and instrumentation
architectures supporting AI and machine learning model-serving operations across
Unclassified, Secret, and Top Secret enclaves within the War Data Platform (WDP)
Core Integration enterprise.
• Develops semantic conventions, runtime instrumentation patterns, and telemetry
pipelines that generate latency metrics, error signatures, throughput
indicators, model-specific performance signals, and operational readiness
measurements for deployed models and serving surfaces.
• Integrates observability capabilities into existing data pipelines,
model-deployment workflows, API access patterns, and serving runtime frameworks
to provide mission-relevant monitoring aligned with Combatant Command and Joint
Staff decision-support needs.
• Configures and validates instrumentation using platforms such as
OpenTelemetry, Prometheus, Grafana, Elastic, Splunk, Amazon CloudWatch, and
service mesh telemetry components to deliver real-time visibility into model
behavior, cross-domain access interactions, and pipeline execution
characteristics.
• Conducts observability readiness reviews, supports test and evaluation gates,
and collaborates with cybersecurity personnel to embed anomaly-detection signals
aligned with Zero Trust and DoW cyber standards.
• Works with serving engineers, pipeline engineers, platform teams, and external
provider integration engineers to maintain observability consistency across
enclaves and resolve domain-specific telemetry constraints.
• Produces observability standards, instrumentation specifications, dashboards,
alerting configurations, and performance analysis reports that strengthen
reliability, accelerate incident response, and reinforce mission assurance for
production model access across all security networks.
• Performs other duties as assigned.
Qualifications
- • Current Secret security clearance with the ability to obtain and maintain a
- Top Secret (TS) security clearance with Sensitive Compartmented Information
- (SCI).
- • 10 or more years of progressive experience in systems engineering, platform
- operations, or ML/AI infrastructure roles, with a demonstrated focus on
- observability, telemetry, and monitoring in classified or federal government
- cloud environments.
- • Hands-on experience designing and implementing observability pipelines using
- industry-standard tooling such as OpenTelemetry, Prometheus, Grafana, Elastic,
- Splunk, or Amazon CloudWatch, including instrumentation of AI/ML model-serving
- runtimes and data pipelines.
- • Experience operating across multi-enclave environments, including NIPRNet,
- SIPRNet, and JWICS, with demonstrated ability to adapt telemetry and
- observability architectures to cross-domain constraints and multi-level security
- requirements.
- • CompTIA Cloud+ certification or equivalent, demonstrating foundational
- knowledge of cloud infrastructure, security, and operational monitoring
- standards.
- • Strong problem-solving and decision-making capabilities, with a proven ability
- to weigh the relative costs and benefits of potential actions and identify the
- most appropriate solution.
- • Highly developed interpersonal and oral/written communication skills, with the
- ability to effectively and professionally interact with a diverse set of
- stakeholders (from peers to end-users to executive management).
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
AI Outcome Customer Engineer, Forward Deployed Engineering
Google · Atlanta, Georgia, United States
Analytics Sr Software Engineer (US Federal)
Workday · Reston, Virginia, United States
AI Evaluations Engineer, US Decision Intelligence
Apple · Cupertino, California, United States
AI Risk Engineer
Bright Vision Technologies · Columbus, Ohio, United States
Research Engineer, Responsible Frontier AI Research, DeepMind
Google · New York, New York, United States
Orchestration Workload Engineer - ACE - AI Factory
Roche · Kaiseraugst, Aargau, Switzerland
Role information can change. Confirm current details on the original application page.
