About the role
What will you do at ARA?
Essential Functions:
* Partner with software developers, platform engineers, and IT staff to improve
system design, operability, deployment safety, and production support
readiness.
* Define and maintain operational standards, runbooks, support procedures,
escalation paths, and service-level objectives.
* Evaluate system architecture and changes to ensure they balance functional
requirements, service quality, reliability, security, and compliance needs.
* Drive continuous improvement in platform stability, maintenance, and
availability.
* Provide advanced technical support and troubleshooting for complex platform
and service issues affecting internal users and stakeholders.
Experience and Skills Required:
* 8+ years of experience in Site Reliability Engineering, DevOps, Platform
Engineering, Systems Engineering, or related infrastructure roles supporting
production services.
* Strong experience with Linux systems administration and troubleshooting in
enterprise environments.
* Strong experience operating and maintaining on-prem Kubernetes platforms and
all related components including CRI, CNI, and CSI plugins.
* Experience deploying and maintaining applications on Kubernetes using Helm,
Kustomize, and similar tooling.
* Experience supporting DevOps tooling such as GitLab, Artifactory, Jira,
Confluence.
* Experience with GitOps tools such as FluxCD or ArgoCD.
* Proficiency scripting with at least one of Python, Go, or Bash.
* Strong experience designing, maintaining, and maturing observability tooling
including monitoring, dashboards, logging and tracing, and supporting SLOs.
* Strong understanding of reliability engineering concepts:
* Service health indicators
* High availability design, failure reduction, and testing
* Operational readiness practices, including developing documentation,
runbooks, and architectural descriptions
* Incident response, root cause analysis, remediation/recovery
* Ability to obtain a security clearance, which includes U.S. citizenship.
Preferred:
* Experience with multiple Linux distributions including Ubuntu.
* Experience with at least one of the following: Tanzu Kubernetes, Nutanix
Kubernetes Platform, Canonical Kubernetes.
* Experience with cloud platforms such as AWS and Azure.
* Experience with infrastructure automation and configuration management.
* Experience managing AI tooling on Kubernetes including MCP Servers, LLM
platforms (vLLM, Ollama), Kubeflow.
* Experience with security and compliance considerations in regulated
environments.
* DoD experience.
* Active or inactive Secret Security Clearance.
Education:
* Bachelor’s degree in CS, Software Engineering or other IT-related field or
equivalent experience
REMOTE WORK NOTICE: This position may be performed fully remote, hybrid, or
onsite at an ARA office. Preference will be given to candidates located onsite
in the Albuquerque area.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
AI Infrastructure AI/ML Engineer (100 % remote) (m/f/d)
EWOR · Capon Bridge, West Virginia, United States
Role information can change. Confirm current details on the original application page.
