Prepin
Log in
JPMorganChase

engineering opportunity

Lead Site Reliability Engineer

The Lead Site Reliability Engineer will guide resiliency design reviews, mentor junior engineers, and implement automated CI/CD pipelines. They are responsible for maintaining 24/7 production support and leveraging AI-assisted workflows to resolve complex technical issues.

Columbus, Ohio, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at JPMorganChase?

Assume a critical role in defining the future of a globally recognized firm and

have a direct and significant effect in a realm tailored for top achievers in

site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise

technology, engineering services and platform team, you hold a leadership role

in your team, demonstrate strong knowledge across multiple technical domains,

and advise others on the technical and business issues facing them. Take lead

and conduct resiliency design reviews, break up complex problems into digestible

work for other engineers, act as a technical lead for medium to large-sized

products, and provide advice and mentoring to other engineers.

Job responsibilities

* Guides and assists others in the areas of building appropriate level designs

and gaining consensus from peers where appropriate

* Collaborates with other software engineers and teams to design and implement

deployment approaches using automated continuous integration and continuous

delivery pipelines

* Collaborates with other software engineers and teams to design, develop,

test, and implement availability, reliability, scalability, and solutions in

their applications

* Implements infrastructure, configuration, and network as code for the

applications and platforms in your remit

* Collaborates with technical experts, key stakeholders, and team members to

resolve complex problems

* Understands service level indicators and utilizes service level objectives to

proactively resolve issues before they impact customers

* Supports the adoption of site reliability engineering best practices within

your team

* Production 24*7 support for business-critical applications

* Uses enterprise-authorized AI capabilities within the work environment to

accelerate major-incident triage, troubleshooting, and post-incident

analysis, validating outputs and handling operational data according to

sensitivity and security requirements.

* Leads reuse-first adoption of AI-assisted reliability workflows across

SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation

automation, and operational readiness), ensuring traceability/auditability,

resiliency, and security controls.

Required qualifications, capabilities, and skills

* Formal training or certification on site reliability engineering concepts and

5+ years applied experience

* Proficient in site reliability engineering (SRE) culture and principles, with

experience implementing SRE practices within applications and platforms;

strong observability background including white/black-box monitoring,

SLO-based alerting, and telemetry collection using tools such as Grafana,

Dynatrace, Prometheus, Datadog, Splunk, and similar.

* Proficient in at least one programming language (e.g., Python, Java/Spring

Boot,. NET) with strong knowledge of software applications and technical

processes within a technical discipline such as cloud, artificial

intelligence, Android, or related areas.

* Demonstrated experience using enterprise-authorized AI capabilities within

the work environment to improve SRE workflows (e.g., incident investigation

support and knowledge capture) with strong validation habits and awareness of

data sensitivity.

* Ability to evaluate AI-assisted operational recommendations for correctness

and risk, define appropriate guardrails for team usage, and ensure outcomes

align to resiliency and security expectations.

* Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and

infrastructure automation using Terraform to build reliable, repeatable

delivery pipelines.

* Strong familiarity with containers and orchestration platforms (Docker,

Kubernetes, ECS), including deploying, scaling, and operating containerized

services in production.

* Proven ability to troubleshoot and resolve common networking issues (DNS,

TCP/IP, routing, TLS, load balancing), applying structured debugging to

restore service quickly.

* Collaborative, proactive team contributor: communicates clearly and

persuasively with minimal supervision, identifies roadblocks early, learns

new technologies quickly, and has experience with event streaming platforms

such as Kafka.

Preferred qualifications, capabilities, and skills

* Ability to identify new technologies and relevant solutions to ensure design

constraints are met by the software team

* Proven track record of initiating and executing ideas that address complex

business challenges

* Deep expertise in networking and systems, including TCP/IP, DNS, load

balancing, firewalls, and VPN technologies; strong Linux performance tuning

and system-level troubleshooting skills

* Certifications a plus: AWS Certified SysOps Administrator or AWS

Professional, Certified Kubernetes Administrator (CKA), Terraform Associate

(or equivalent)

* Collaborative leader with a proven track record mentoring junior engineers,

driving SRE best-practice adoption across teams, and communicating clearly to

both technical and non-technical stakeholders (including presentations)

* Experience in handling critical incident and change management – be part of

critical incident taskforce call.

* Familiarity of agile practices – preferably, scrum and Kanban

JPMorganChase, one of the oldest financial institutions, offers innovative

financial solutions to millions of consumers, small businesses and many of the

world’s most prominent corporate, institutional and government clients under the

J.P. Morgan and Chase brands. Our history spans over 200 years and today we are

a leader in investment banking, consumer and small business banking, commercial

banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined

based on the role, experience, skill set and location. Those in eligible roles

may receive commission-based pay and/or discretionary incentive compensation,

paid in the form of cash and/or forfeitable equity, awarded in recognition of

individual achievements and contributions. We also offer a range of benefits and

programs to meet employee needs, based on eligibility. These benefits include

comprehensive health care coverage, on-site health and wellness centers, a

retirement savings plan, backup childcare, tuition reimbursement, mental health

support, financial coaching and more. Additional details about total

compensation and benefits will be provided during the hiring process.

We recognize that our people are our strength and the diverse talents they bring

to our global workforce are directly linked to our success. We are an equal

opportunity employer and place a high value on diversity and inclusion at our

company. We do not discriminate on the basis of any protected attribute,

including race, religion, color, national origin, gender, sexual orientation,

gender identity, gender expression, age, marital or veteran status, pregnancy or

disability, or any other basis protected under applicable law. We also make

reasonable accommodations for applicants’ and employees’ religious practices and

beliefs, as well as mental health or physical disability needs. Visit our FAQs

[https://careers.jpmorgan.com/us/en/how-we-hire/faqs] for more information about

requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including

Disability/Veterans

Our professionals in our Corporate Functions cover a diverse range of areas from

finance and risk to human resources and marketing. Our corporate teams are an

essential part of our company, ensuring that we’re setting our businesses,

clients, customers and employees up for success.

Which skills does this role require?

Site Reliability EngineeringObservabilityPythonJavaCI/CDDockerKafkaCloudArtificial IntelligenceSystem troubleshootingSpring BootJenkinsGitLabGrafanaDynatracePrometheusDatadogSplunkResiliencyAutomationIncident ManagementC#

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.