Prepin
Log in
JPMorganChase

engineering opportunity

Lead Site Reliability Engineer - Network

The Lead Site Reliability Engineer will oversee the reliability and performance of mission-critical network services while mentoring engineering teams. They are responsible for leading major incident responses, conducting resiliency design reviews, and driving automation to reduce toil.

Columbus, Ohio, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at JPMorganChase?

Assume a critical role in defining the future of a globally recognized firm and

have a direct and significant effect in a realm tailored for top achievers in

site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise

technology, Infrastructure Platforns team, you hold a leadership role in your

team, demonstrate strong knowledge across multiple technical domains, and advise

others on the technical and business issues facing them. Take lead and conduct

resiliency design reviews, break up complex problems into digestible work for

other engineers, act as a technical lead for medium to large-sized products, and

provide advice and mentoring to other engineers.

Job responsibilities

* Consistently models and champions site reliability culture and practices,

documents and shares knowledge within your organization via internal forums

and communities of practice

* Leads initiatives to improve the reliability and stability of your team’s

applications and platforms using data-driven analytics to improve service

levels, proactively identifying and solving technology-related bottlenecks in

areas of expertise

* Drives collaboration with your team to identify comprehensive service level

indicators and the stakeholder partners to establish reasonable service level

objectives and error budgets with your customers

* Uses enterprise-authorized AI capabilities within the work environment to

accelerate major-incident triage, troubleshooting, and post-incident

analysis, validating outputs and handling operational data according to

sensitivity and security requirements.

* Serves as the main point of contact during major incidents for your

application and has the skills to identify and solve the issue quickly to

avoid financial loss to the business

* Offers a high level of technical expertise within one or more technical

domains and provides advice and mentorship to other engineers

* Own reliability for mission-critical network services (availability,

performance, recoverability), leading major incident response, and driving

strong post-incident learning.

* Own problem management end-to-end: author/approve RCAs, identify systemic

root causes, and deliver durable remediation that prevents repeat incidents.

* Architect and deliver automation using Python, Shell, and Ansible

(self-healing where appropriate, automated remediation with guardrails,

validation pipelines, and standardized tooling) to reduce toil and improve

reliability.

* Provide deep technical leadership across SD-WAN/SDA/SDN, routing & switching,

and network security/traffic services (firewalls, load balancers, proxies),

partnering with platform teams on observability and embedding SRE practices

(NFR targets, FMEA-style risk analysis, readiness reviews, resilience

testing, and change safety).

* Leads reuse-first adoption of AI-assisted reliability workflows across

SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation

automation, and operational readiness), ensuring traceability/auditability,

resiliency, and security controls.

Required qualifications, capabilities, and skills

* Formal training or certification on site reliability engineering concepts and

5+ years applied experience

* Demonstrated proficiency in reliability, scalability, performance, security,

enterprise system architecture, toil reduction, and other site reliability

best practices

* Fluent in at least one programming language such as: Python, Java/Spring

Boot, . Net

* Demonstrated experience using enterprise-authorized AI capabilities within

the work environment to improve SRE workflows (e.g., incident investigation

support and knowledge capture) with strong validation habits and awareness of

data sensitivity.

* Ability to evaluate AI-assisted operational recommendations for correctness

and risk, define appropriate guardrails for team usage, and ensure outcomes

align to resiliency and security expectations.

* Proficient knowledge and experience in observability such as white and black

box monitoring, service level objective alerting, and telemetry collection

* Proficient with continuous integration and continuous delivery practices and

tooling

* Proficient with container and container orchestration

* Experience with troubleshooting common networking technologies and issues

* Advanced knowledge of software applications and technical processes with

emerging depth in one or more technical disciplines, and actively

self-educates to evaluate and recommend suitable new technologies

Preferred qualifications, capabilities, and skills

* Extensive experience operating and engineering large-scale networks with

strong troubleshooting depth.

* Proven leadership in major incidents, RCA, and delivery of long-term fixes

that reduce repeat incidents.

* Advanced automation track record using Python, Shell, and Ansible, with

demonstrated toil reduction and reliability gains.

* Strong applied understanding of SRE concepts, NFRs, and FMEA (or equivalent

failure/risk analysis methods).

* Strong ownership mindset, stakeholder management, and ability to drive

delivery independently.

JPMorganChase, one of the oldest financial institutions, offers innovative

financial solutions to millions of consumers, small businesses and many of the

world’s most prominent corporate, institutional and government clients under the

J.P. Morgan and Chase brands. Our history spans over 200 years and today we are

a leader in investment banking, consumer and small business banking, commercial

banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined

based on the role, experience, skill set and location. Those in eligible roles

may receive commission-based pay and/or discretionary incentive compensation,

paid in the form of cash and/or forfeitable equity, awarded in recognition of

individual achievements and contributions. We also offer a range of benefits and

programs to meet employee needs, based on eligibility. These benefits include

comprehensive health care coverage, on-site health and wellness centers, a

retirement savings plan, backup childcare, tuition reimbursement, mental health

support, financial coaching and more. Additional details about total

compensation and benefits will be provided during the hiring process.

We recognize that our people are our strength and the diverse talents they bring

to our global workforce are directly linked to our success. We are an equal

opportunity employer and place a high value on diversity and inclusion at our

company. We do not discriminate on the basis of any protected attribute,

including race, religion, color, national origin, gender, sexual orientation,

gender identity, gender expression, age, marital or veteran status, pregnancy or

disability, or any other basis protected under applicable law. We also make

reasonable accommodations for applicants’ and employees’ religious practices and

beliefs, as well as mental health or physical disability needs. Visit our FAQs

[https://careers.jpmorgan.com/us/en/how-we-hire/faqs] for more information about

requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including

Disability/Veterans

Our professionals in our Corporate Functions cover a diverse range of areas from

finance and risk to human resources and marketing. Our corporate teams are an

essential part of our company, ensuring that we’re setting our businesses,

clients, customers and employees up for success.

Which skills does this role require?

Site Reliability EngineeringJavaSpring BootSD-WANSDNNetwork SecurityObservabilityCI/CDContainer OrchestrationIncident ResponseRoot Cause AnalysisSystem ArchitectureNetwork EngineeringSDAFirewallsLoad BalancersProxiesIncident ManagementNFRInfrastructure PlatformsEnterprise ArchitectureAI-assisted workflowsTelemetryResiliencyScalabilitySecurityService Level ObjectivesC#

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.