About the role
What will you do at JPMorganChase?
Assume a critical role in defining the future of a globally recognized firm and
have a direct and significant effect in a realm tailored for top achievers in
site reliability.
As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise
technology, Infrastructure Platforns team, you hold a leadership role in your
team, demonstrate strong knowledge across multiple technical domains, and advise
others on the technical and business issues facing them. Take lead and conduct
resiliency design reviews, break up complex problems into digestible work for
other engineers, act as a technical lead for medium to large-sized products, and
provide advice and mentoring to other engineers.
Job responsibilities
* Consistently models and champions site reliability culture and practices,
documents and shares knowledge within your organization via internal forums
and communities of practice
* Leads initiatives to improve the reliability and stability of your team’s
applications and platforms using data-driven analytics to improve service
levels, proactively identifying and solving technology-related bottlenecks in
areas of expertise
* Drives collaboration with your team to identify comprehensive service level
indicators and the stakeholder partners to establish reasonable service level
objectives and error budgets with your customers
* Uses enterprise-authorized AI capabilities within the work environment to
accelerate major-incident triage, troubleshooting, and post-incident
analysis, validating outputs and handling operational data according to
sensitivity and security requirements.
* Serves as the main point of contact during major incidents for your
application and has the skills to identify and solve the issue quickly to
avoid financial loss to the business
* Offers a high level of technical expertise within one or more technical
domains and provides advice and mentorship to other engineers
* Own reliability for mission-critical network services (availability,
performance, recoverability), leading major incident response, and driving
strong post-incident learning.
* Own problem management end-to-end: author/approve RCAs, identify systemic
root causes, and deliver durable remediation that prevents repeat incidents.
* Architect and deliver automation using Python, Shell, and Ansible
(self-healing where appropriate, automated remediation with guardrails,
validation pipelines, and standardized tooling) to reduce toil and improve
reliability.
* Provide deep technical leadership across SD-WAN/SDA/SDN, routing & switching,
and network security/traffic services (firewalls, load balancers, proxies),
partnering with platform teams on observability and embedding SRE practices
(NFR targets, FMEA-style risk analysis, readiness reviews, resilience
testing, and change safety).
* Leads reuse-first adoption of AI-assisted reliability workflows across
SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation
automation, and operational readiness), ensuring traceability/auditability,
resiliency, and security controls.
Required qualifications, capabilities, and skills
* Formal training or certification on site reliability engineering concepts and
5+ years applied experience
* Demonstrated proficiency in reliability, scalability, performance, security,
enterprise system architecture, toil reduction, and other site reliability
best practices
* Fluent in at least one programming language such as: Python, Java/Spring
Boot, . Net
* Demonstrated experience using enterprise-authorized AI capabilities within
the work environment to improve SRE workflows (e.g., incident investigation
support and knowledge capture) with strong validation habits and awareness of
data sensitivity.
* Ability to evaluate AI-assisted operational recommendations for correctness
and risk, define appropriate guardrails for team usage, and ensure outcomes
align to resiliency and security expectations.
* Proficient knowledge and experience in observability such as white and black
box monitoring, service level objective alerting, and telemetry collection
* Proficient with continuous integration and continuous delivery practices and
tooling
* Proficient with container and container orchestration
* Experience with troubleshooting common networking technologies and issues
* Advanced knowledge of software applications and technical processes with
emerging depth in one or more technical disciplines, and actively
self-educates to evaluate and recommend suitable new technologies
Preferred qualifications, capabilities, and skills
* Extensive experience operating and engineering large-scale networks with
strong troubleshooting depth.
* Proven leadership in major incidents, RCA, and delivery of long-term fixes
that reduce repeat incidents.
* Advanced automation track record using Python, Shell, and Ansible, with
demonstrated toil reduction and reliability gains.
* Strong applied understanding of SRE concepts, NFRs, and FMEA (or equivalent
failure/risk analysis methods).
* Strong ownership mindset, stakeholder management, and ability to drive
delivery independently.
JPMorganChase, one of the oldest financial institutions, offers innovative
financial solutions to millions of consumers, small businesses and many of the
world’s most prominent corporate, institutional and government clients under the
J.P. Morgan and Chase brands. Our history spans over 200 years and today we are
a leader in investment banking, consumer and small business banking, commercial
banking, financial transaction processing and asset management.
We offer a competitive total rewards package including base salary determined
based on the role, experience, skill set and location. Those in eligible roles
may receive commission-based pay and/or discretionary incentive compensation,
paid in the form of cash and/or forfeitable equity, awarded in recognition of
individual achievements and contributions. We also offer a range of benefits and
programs to meet employee needs, based on eligibility. These benefits include
comprehensive health care coverage, on-site health and wellness centers, a
retirement savings plan, backup childcare, tuition reimbursement, mental health
support, financial coaching and more. Additional details about total
compensation and benefits will be provided during the hiring process.
We recognize that our people are our strength and the diverse talents they bring
to our global workforce are directly linked to our success. We are an equal
opportunity employer and place a high value on diversity and inclusion at our
company. We do not discriminate on the basis of any protected attribute,
including race, religion, color, national origin, gender, sexual orientation,
gender identity, gender expression, age, marital or veteran status, pregnancy or
disability, or any other basis protected under applicable law. We also make
reasonable accommodations for applicants’ and employees’ religious practices and
beliefs, as well as mental health or physical disability needs. Visit our FAQs
[https://careers.jpmorgan.com/us/en/how-we-hire/faqs] for more information about
requesting an accommodation.
JPMorgan Chase & Co. is an Equal Opportunity Employer, including
Disability/Veterans
Our professionals in our Corporate Functions cover a diverse range of areas from
finance and risk to human resources and marketing. Our corporate teams are an
essential part of our company, ensuring that we’re setting our businesses,
clients, customers and employees up for success.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Data Center Engineer – Windows Infrastructure
Konnect IT Group, Inc. · Chicago, Illinois, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
AI Infrastructure AI/ML Engineer (100 % remote) (m/f/d)
EWOR · Capon Bridge, West Virginia, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Role information can change. Confirm current details on the original application page.
