Prepin
Log in
ADT Security Services

engineering opportunity

Senior Site Reliability Engineer

The Senior Site Reliability Engineer will own the reliability, scalability, and performance of production platforms while driving operational excellence through automation and observability. They will also lead incident response efforts, mentor junior engineers, and collaborate with cross-functional teams to improve platform stability.

Whitpain Township, Pennsylvania, United StateshybridFULL_TIME

Posted

About the role

What will you do at ADT Security Services?

Applicants must be authorized to work for any employer in the United States

without current or future sponsorship. ADT is unable to provide employment-based

immigration sponsorship, including but not limited to H-1B, TN, F-1 STEM OPT, or

other work authorization sponsorship.

Locations and Workstyle:

Blue Bell, PA: Primarily remote; candidates should be within commuting distance

of the Blue Bell office and able to work onsite as needed. Option to come onsite

more frequently if desired.

Irving, TX and Boca Raton, FL: Hybrid schedule - onsite a minimum of four days

per week, with one remote day. Five days onsite may be required based on

business needs.

What You'll Do

  • * Own the reliability, availability, scalability, and performance of major
  • production platforms and services.
  • * Work closely with Infrastructure, Development, Product, and Operations teams
  • to keep the ADT platform running and customers protected.
  • * Drive operational excellence through automation, observability, and proactive
  • problem-solving across large-scale distributed systems.
  • * Lead reliability efforts across cloud environments, leveraging tools such as
  • Terraform, Ansible, Kubernetes, Dynatrace, and Prometheus.
  • * Design, build, and evolve infrastructure as code solutions to automate
  • provisioning, configuration management, patching, and repeatable operational
  • processes.
  • * Own the operational lifecycle of critical services, including production
  • readiness, monitoring strategy, incident response, and continuous
  • improvement.
  • * Identify reliability gaps, performance bottlenecks, and operational toil,
  • then implement durable solutions that improve stability and efficiency.
  • * Define and improve observability practices, including dashboards, alerts,
  • runbooks, and detection strategies that reduce time to detect and resolve
  • issues.
  • * Support software releases and production changes, including validation,
  • rollback planning, post-change verification, and operational readiness
  • reviews.
  • * Partner with cross-functional teams to improve operational health and
  • establish engineering practices that strengthen platform reliability.
  • * Mentor junior and mid-level engineers through design reviews, technical
  • guidance, on-call coaching, and operational best practices.
  • * Participate in an on-call rotation and provide production support, including
  • complex incident response, root cause analysis, remediation efforts, and
  • support for customer-impacting issues during major incidents.
  • What You'll Need
  • * 7+ years of experience in Site Reliability Engineering, Infrastructure
  • Engineering, Systems Engineering,
  • * Operations Engineering, DevOps, or related production-focused roles with
  • on-call responsibility.
  • * Proven experience owning large-scale distributed systems and platforms in
  • production environments.
  • * Strong Linux and systems engineering fundamentals, including troubleshooting
  • across compute, storage, networking, and application layers.
  • * Advanced experience with infrastructure as code, including Terraform and
  • Ansible design, implementation, and maintenance.
  • * Experience operating and optimizing cloud environments, including AWS and
  • GCP.
  • * Strong experience managing Kubernetes clusters in large-scale production
  • environments.
  • * Proficiency in Python, Bash, or similar scripting languages, with a focus on
  • automation and operational tooling.
  • * Strong understanding of software delivery, change management, and safe
  • production operations.
  • * Experience with monitoring and observability platforms such as Dynatrace,
  • Prometheus, or similar technologies.
  • * Ability to diagnose and resolve complex production issues while making sound
  • decisions around risk, rollback strategies, and escalation paths.
  • * Experience with CI/CD pipelines and deployment automation.
  • * Experience leading or supporting incident response activities and
  • post-incident reviews.
  • * Strong communication skills and the ability to collaborate effectively across
  • technical and business teams.
  • * Comfortable operating in complex environments with ambiguity and competing
  • priorities.
  • * Ability to balance incident-driven operational work with longer-term
  • reliability and automation initiatives.
  • * Comfortable using AI tools and agents to accelerate investigation,
  • automation, and documentation while maintaining sound engineering judgment.

Preferred Qualifications

  • * Experience with Java/JVM ecosystems and large-scale customer-facing
  • platforms.
  • * Experience with Kafka or other distributed messaging platforms.
  • * Experience leading security remediation efforts at scale, including patch
  • SLAs, CVE response, and operating system upgrades.
  • * Familiarity with Harness, enterprise Git workflows, and audit-driven change
  • controls.
  • * Demonstrated success improving reliability metrics such as MTTD, MTTR,
  • availability, or operational efficiency across teams.

Compensation

&

Benefits

The salary range for this role is $128,800.00 - $193,200.00 and is based on

experience and qualifications.

Certain roles are eligible for annual bonus and may include equity. These awards

are allocated based on company and individual performance.

We offer employees access to healthcare benefits, a 401(k) plan and company

match, short-term and long-term disability coverage, life insurance, wellbeing

benefits and paid time off among others. Employees accrue up to 120 hours in

their first year. Your accrual rate increases after your first year. We also

offer 6 paid holidays.

Which skills does this role require?

Site Reliability EngineeringInfrastructure as CodeKubernetesTerraformAnsibleLinuxPythonBashAWSGCPDynatracePrometheusCI/CDIncident ResponseObservabilityDistributed SystemsAutomationMonitoringProduction ReadinessRoot Cause AnalysisScalabilityPatch ManagementOperations

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.