Prepin
Log in
Goldman Sachs

engineering opportunity

Site Reliability Engineer, Global Banking & Markets, Vice President

The Site Reliability Engineer will ensure the availability, resilience, and performance of core trading services by defining SLOs and automating operational toil. They will lead incident response efforts and orchestrate AI-driven tools to accelerate root-cause analysis and system modernization.

New York, New York, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Goldman Sachs?

What We Do

At Goldman Sachs, our Engineers don't just make things - we make things

possible. Change the world by connecting people and capital with ideas. Solve

the most challenging and pressing engineering problems for our clients. Join

our engineering teams that build massively scalable software and systems,

architect low latency infrastructure solutions, proactively guard against cyber

threats, and leverage machine learning alongside financial engineering to

continuously turn data into action. Create new businesses, transform finance,

and explore a world of opportunity at the speed of markets.

Within the firm's Global Banking & Markets business, the Site Reliability

Engineering (SRE) team ensures the availability, resilience, and performance of

core business services that underpin a global 24×7 trading operation. Working

across Global Markets' front, middle, and back office functions, you will

engineer reliability while balancing stringent non-functional demands

for availability, latency, and resilience as well as complex, evolving business

requirements.

Want to push the limit of digital possibilities? Start here.

Who We Look For

Goldman Sachs Engineers are at the forefront of innovation, driving solutions as

creative collaborators in a fast-paced global environment. We seek individuals

who evolve, adapt, and thrive on challenging problems.

As part of our SRE team, you will operate at the intersection of reliability

engineering, cloud infrastructure, and AI-driven operations. Using Goldman

Sachs' AI tooling and agentic assistants, you will accelerate incident

diagnosis, automate operational toil, comprehend large legacy codebases, and

raise the bar for production-quality automation across the software and

reliability lifecycle. Above all, you will bring strong risk acumen and the

ability to connect the right people across the organization to resolve problems

quickly and decisively.

Your Impact

* Own reliability outcomes: Define and defend Service Level Objectives (SLOs),

error budgets, and reliability standards for critical trading services, with

risk always front of mind.

* Reduce risk and toil: Identify systemic risks before they materialize,

automate away repetitive operational work, and strengthen the resilience

posture of the platform.

* Connect and communicate: Act as a trusted coordinator during incidents —

rapidly mobilizing the right engineers, domain experts, and stakeholders

across a globally distributed organization, and communicating clearly with

both technical and non-technical audiences.

* Multiply your output with AI: Orchestrate AI coding and operations agents to

accelerate root-cause analysis, remediation, and automation while maintaining

mastery, quality, and production fitness over all AI-generated work.

* Build for the future: Design and operate high-availability, multi-region,

event-driven services on a modern cloud-native platform, setting the

reliability and architectural standard for years to come.

What You Will Do

  • * Design, build, and operate high-availability, multi-region, cloud-native
  • services with security and comprehensive observability (metrics, distributed
  • tracing, structured logging) built in at every layer.
  • * Establish and manage SLIs, SLOs, and error budgets; drive blameless
  • post-incident reviews and translate findings into durable engineering
  • improvements.
  • * Lead incident response for latency-sensitive, high-throughput trade lifecycle
  • systems — quickly diagnosing issues, coordinating cross-functional
  • responders, and communicating status to stakeholders.
  • * Develop event-driven architectures, multi-stage processing pipelines, and
  • optimized data paths for high-throughput trade lifecycle management.
  • * Apply strong risk acumen to change management, capacity planning, and
  • resilience testing (chaos engineering, failover, and BCP drills).
  • * Partner with engineers, domain experts, and global stakeholders to understand
  • production processes, challenge entrenched assumptions in a cloud-centric,
  • AI-driven world, and drive modernization.
  • * Multiply your impact with a modern, AI-centric toolchain, orchestrating AI
  • agents across the SDLC and operations to rapidly comprehend large codebases,
  • generate production-quality automation, and accelerate delivery.
  • Basic Qualifications
  • * 8+ years of professional software / reliability engineering experience, with
  • strong command of at least one major language (Java 17+ preferred), including
  • concurrency, collections, and modern language features.
  • * Demonstrated risk acumen — the ability to identify, quantify, and mitigate
  • operational and technical risk in a regulated financial services environment.
  • * Excellent communication and stakeholder-coordination skills — proven ability
  • to connect the right people quickly and drive resolution across
  • geographically distributed, technical and non-technical audiences.
  • * Proven experience running high-availability production environments:
  • SLIs/SLOs, error budgets, on-call, incident command, and post-incident
  • reviews.
  • * Strong understanding of cloud infrastructure (GCP, AWS), container
  • orchestration (Kubernetes, Docker), and infrastructure-as-code.
  • * Working knowledge of AI models and AI-assisted engineering tools (e.g.,
  • Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist), including
  • the ability to govern AI agents, critically assess their output, and maintain
  • quality over AI-generated work.
  • * Experience building event-driven and distributed systems, including messaging
  • platforms (e.g., Apache Kafka), delivery guarantees, and resilience
  • strategies.
  • * Strong SDLC and automation practices: version control, CI/CD pipelines,
  • automated build/test/deploy workflows, and code quality tooling.
  • * Solid observability discipline: application instrumentation, distributed
  • tracing, structured logging, and metrics-driven operations.
  • * Ability to rapidly navigate, understand, and debug large and unfamiliar
  • codebases — with and without AI assistance.

Preferred Qualifications

  • Experience with a meaningful subset of the following is highly valued:
  • * Reliability & Operations: Chaos engineering, capacity planning,
  • load/performance testing, and production support in high-availability,
  • latency-sensitive environments.
  • * Frameworks & Architecture: Spring Boot, gRPC / Protocol Buffers,
  • integration/orchestration frameworks (e.g., Apache Camel, Spring
  • Integration), and pipeline/adapter patterns (retry, dead-letter queues, error
  • isolation).
  • * Cloud & Infrastructure: Cloud platforms (GCP, AWS), Kubernetes/Docker, JVM
  • tuning for containerized workloads, and infrastructure-as-code (Terraform,
  • Helm).
  • * AI & Automation: Applying AI models to operational use cases — anomaly
  • detection, log analysis, automated remediation, and agentic operations.
  • * Observability & Operations: Prometheus, Grafana, OpenTelemetry, and SLO
  • tooling.
  • * Data & Performance: Data modeling, SQL/NoSQL databases, caching strategies,
  • and performance optimization in latency-sensitive systems.
  • * Security: Enterprise security patterns; authentication protocols, mutual TLS,
  • secrets management, and certificate rotation.
  • * Domain Knowledge: Equities, post-trade, or financial services experience;
  • trade lifecycle concepts, position management, reconciliation, and
  • multi-system migration environments.
  • * Other: Asynchronous / non-blocking I/O frameworks (e.g., Vert.x, Netty),
  • multi-region / BCP architectures, and open-source contribution experience.
  • Salary Range
  • The expected base salary for this New York, New York, United States-based
  • position is $150,000-$300,000. In addition, you may be eligible for a
  • discretionary bonus if you are an active employee as of fiscal year-end.

Benefits

Goldman Sachs is committed to providing our people with valuable and competitive

benefits and wellness offerings, as it is a core part of providing a strong

overall employee experience. A summary of these offerings, which are generally

available to active, non-temporary, full-time and part-time US employees who

work at least 20 hours per week, can be found here

[https://www.goldmansachs.com/careers/discover/2022-Benefits-Summary-US.pdf].

ABOUT GOLDMAN SACHS

At Goldman Sachs, we commit our people, capital and ideas to help our clients,

shareholders and the communities we serve to grow. Founded in 1869, we are a

leading global investment banking, securities and investment management firm.

Headquartered in New York, we maintain offices around the world.

We believe who you are makes you better at what you do. We're committed to

fostering and advancing diversity and inclusion in our own workplace and beyond

by ensuring every individual within our firm has a number of opportunities to

grow professionally and personally, from our training and development

opportunities and firmwide networks to benefits, wellness and personal finance

offerings and mindfulness programs. Learn more about our culture, benefits, and

people at GS.com/careers.

We’re committed to finding reasonable accommodations for candidates with special

needs or disabilities during our recruiting process. Learn more:

https://www.goldmansachs.com/careers/footer/disability-statement.html

© The Goldman Sachs Group, Inc., 2026. All rights reserved.

Which skills does this role require?

Site Reliability EngineeringJavaIncident ResponseRisk managementChaos engineeringPerformance optimizationSystem architectureTerraformHelmPrometheusGrafanaOpenTelemetrySpring BootgRPCDistributed tracingTrade lifecycleLatency-sensitive systemsEvent-driven architectureSQLNoSQLProduction supportCapacity planningSecurityMachine Learning

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.