Prepin
Log in
Symphony

engineering opportunity

Senior Site Reliability Engineer

Design and maintain scalable cloud infrastructure using IaC and GitOps principles while managing Kubernetes clusters. Provide production support and implement observability practices to ensure high availability and system reliability.

New York, New York, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Symphony?

About us @Symphony Secure. Connected. Intelligent.

Symphony is an AI-powered communication and technology company fueled by interconnected platforms: messaging, voice, directory and analytics. Our end-to-end encrypted technologies enable over 1,400 institutions to accelerate AI impact, prioritize data security, navigate complex regulatory compliance and optimize business interactions. Role

Description

As a Site Reliability Engineer, you will work with Agile engineering teams to provide production insight into running and operating software at-scale in a globally distributed and highly available cloud based system. You will guide the team to consider resiliency, scalability and operability implications in the choices they make during the development cycle to help foster an ownership in production mentality ("You build it, you own it").

You will also own and develop our platform by implementing and championing GitOps principles. You will play a hands-on role helping the team meet technical, operational, schedule, and business requirements.

The ideal candidate will be a systems problem solver with a passion for crafting products that deliver incredible customer experiences, have deep experience with infrastructure, operational automation, data driven metrics collection, modern platform management, and a true desire to automate it rather than do it repeatedly.

Key Responsibilities

  • * Design, build, and maintain highly scalable and reliable infrastructure using Infrastructure-as-Code (IaC).
  • * Manage and optimize Kubernetes clusters, leveraging Helm and ArgoCD for efficient application deployment and lifecycle management.
  • * Administer and troubleshoot Linux-based systems, ensuring their performance, security,and availability.
  • * Work extensively with both GCP and AWS services, architecting and managing cloud-native solutions.
  • * Implement robust observability practices, including monitoring, logging and alerting to proactively identify and resolve issues, using Splunk and Grafana.
  • * Develop and maintain automation tools using languages such as Python and Go to streamline operations and improve efficiency.
  • * Provide production support, responding to incidents and outages, and participating in a 24/7 on-call rotation.
  • * Troubleshoot complex networking issues and optimize network performance.
  • * Engage in communications across all areas of the organization.
  • Required

Qualifications

  • * Strong experience with IaC and CMaC, Terraform (nice to have Terragrunt and Ansible) * Deep understanding of Kubernetes and Helm.
  • * Extensive experience with Linux administration and troubleshooting.
  • * Hands-on experience with GCP and AWS services and cloud-native architectures.
  • * Experience in implementing observability solutions.
  • * Solid understanding of networking concepts and protocols.
  • * Ability to work both independently and collaboratively.
  • * Ability to lead and work on projects.
  • * Ability to multitask and adapt quickly to changing priorities.
  • * Excellent communication and problem-solving skills.
  • We have an international team, so good communication skills are essential.
  • * Willingness to participate in a 24/7 on-call rotation.

Compensation

* Salary Range: $140,000 -170,000 base salary per year * Bonus Plan Benefits and Perks: * Regional specific competitive benefits * Build your own Benefits (BYOB) perk * Local events, team building, and development opportunities We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, colour, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment.

Which skills does this role require?

Infrastructure-as-CodeArgoCDSplunkGrafanaPythonGoTerraformGitOpsSRESite Reliability EngineeringAgileScalabilityResiliencyCMaC

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.