About the role
What will you do at Goldman Sachs?
What We Do
At Goldman Sachs, our Engineers don't just make things - we make things
possible. Change the world by connecting people and capital with ideas. Solve
the most challenging and pressing engineering problems for our clients. Join
our engineering teams that build massively scalable software and systems,
architect low latency infrastructure solutions, proactively guard against cyber
threats, and leverage machine learning alongside financial engineering to
continuously turn data into action. Create new businesses, transform finance,
and explore a world of opportunity at the speed of markets.
Within the firm's Global Banking & Markets business, the Site Reliability
Engineering (SRE) team ensures the availability, resilience, and performance of
core business services that underpin a global 24×7 trading operation. Working
across Global Markets' front, middle, and back office functions, you will
engineer reliability while balancing stringent non-functional demands
for availability, latency, and resilience as well as complex, evolving business
requirements.
Want to push the limit of digital possibilities? Start here.
Who We Look For
Goldman Sachs Engineers are at the forefront of innovation, driving solutions as
creative collaborators in a fast-paced global environment. We seek individuals
who evolve, adapt, and thrive on challenging problems.
As part of our SRE team, you will operate at the intersection of reliability
engineering, cloud infrastructure, and AI-driven operations. Using Goldman
Sachs' AI tooling and agentic assistants, you will accelerate incident
diagnosis, automate operational toil, comprehend large legacy codebases, and
raise the bar for production-quality automation across the software and
reliability lifecycle. Above all, you will bring strong risk acumen and the
ability to connect the right people across the organization to resolve problems
quickly and decisively.
Your Impact
* Own reliability outcomes: Define and defend Service Level Objectives (SLOs),
error budgets, and reliability standards for critical trading services, with
risk always front of mind.
* Reduce risk and toil: Identify systemic risks before they materialize,
automate away repetitive operational work, and strengthen the resilience
posture of the platform.
* Connect and communicate: Act as a trusted coordinator during incidents —
rapidly mobilizing the right engineers, domain experts, and stakeholders
across a globally distributed organization, and communicating clearly with
both technical and non-technical audiences.
* Multiply your output with AI: Orchestrate AI coding and operations agents to
accelerate root-cause analysis, remediation, and automation while maintaining
mastery, quality, and production fitness over all AI-generated work.
* Build for the future: Design and operate high-availability, multi-region,
event-driven services on a modern cloud-native platform, setting the
reliability and architectural standard for years to come.
What You Will Do
- * Design, build, and operate high-availability, multi-region, cloud-native
- services with security and comprehensive observability (metrics, distributed
- tracing, structured logging) built in at every layer.
- * Establish and manage SLIs, SLOs, and error budgets; drive blameless
- post-incident reviews and translate findings into durable engineering
- improvements.
- * Lead incident response for latency-sensitive, high-throughput trade lifecycle
- systems — quickly diagnosing issues, coordinating cross-functional
- responders, and communicating status to stakeholders.
- * Develop event-driven architectures, multi-stage processing pipelines, and
- optimized data paths for high-throughput trade lifecycle management.
- * Apply strong risk acumen to change management, capacity planning, and
- resilience testing (chaos engineering, failover, and BCP drills).
- * Partner with engineers, domain experts, and global stakeholders to understand
- production processes, challenge entrenched assumptions in a cloud-centric,
- AI-driven world, and drive modernization.
- * Multiply your impact with a modern, AI-centric toolchain, orchestrating AI
- agents across the SDLC and operations to rapidly comprehend large codebases,
- generate production-quality automation, and accelerate delivery.
- Basic Qualifications
- * 8+ years of professional software / reliability engineering experience, with
- strong command of at least one major language (Java 17+ preferred), including
- concurrency, collections, and modern language features.
- * Demonstrated risk acumen — the ability to identify, quantify, and mitigate
- operational and technical risk in a regulated financial services environment.
- * Excellent communication and stakeholder-coordination skills — proven ability
- to connect the right people quickly and drive resolution across
- geographically distributed, technical and non-technical audiences.
- * Proven experience running high-availability production environments:
- SLIs/SLOs, error budgets, on-call, incident command, and post-incident
- reviews.
- * Strong understanding of cloud infrastructure (GCP, AWS), container
- orchestration (Kubernetes, Docker), and infrastructure-as-code.
- * Working knowledge of AI models and AI-assisted engineering tools (e.g.,
- Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist), including
- the ability to govern AI agents, critically assess their output, and maintain
- quality over AI-generated work.
- * Experience building event-driven and distributed systems, including messaging
- platforms (e.g., Apache Kafka), delivery guarantees, and resilience
- strategies.
- * Strong SDLC and automation practices: version control, CI/CD pipelines,
- automated build/test/deploy workflows, and code quality tooling.
- * Solid observability discipline: application instrumentation, distributed
- tracing, structured logging, and metrics-driven operations.
- * Ability to rapidly navigate, understand, and debug large and unfamiliar
- codebases — with and without AI assistance.
Preferred Qualifications
- Experience with a meaningful subset of the following is highly valued:
- * Reliability & Operations: Chaos engineering, capacity planning,
- load/performance testing, and production support in high-availability,
- latency-sensitive environments.
- * Frameworks & Architecture: Spring Boot, gRPC / Protocol Buffers,
- integration/orchestration frameworks (e.g., Apache Camel, Spring
- Integration), and pipeline/adapter patterns (retry, dead-letter queues, error
- isolation).
- * Cloud & Infrastructure: Cloud platforms (GCP, AWS), Kubernetes/Docker, JVM
- tuning for containerized workloads, and infrastructure-as-code (Terraform,
- Helm).
- * AI & Automation: Applying AI models to operational use cases — anomaly
- detection, log analysis, automated remediation, and agentic operations.
- * Observability & Operations: Prometheus, Grafana, OpenTelemetry, and SLO
- tooling.
- * Data & Performance: Data modeling, SQL/NoSQL databases, caching strategies,
- and performance optimization in latency-sensitive systems.
- * Security: Enterprise security patterns; authentication protocols, mutual TLS,
- secrets management, and certificate rotation.
- * Domain Knowledge: Equities, post-trade, or financial services experience;
- trade lifecycle concepts, position management, reconciliation, and
- multi-system migration environments.
- * Other: Asynchronous / non-blocking I/O frameworks (e.g., Vert.x, Netty),
- multi-region / BCP architectures, and open-source contribution experience.
- Salary Range
- The expected base salary for this New York, New York, United States-based
- position is $150,000-$300,000. In addition, you may be eligible for a
- discretionary bonus if you are an active employee as of fiscal year-end.
Benefits
Goldman Sachs is committed to providing our people with valuable and competitive
benefits and wellness offerings, as it is a core part of providing a strong
overall employee experience. A summary of these offerings, which are generally
available to active, non-temporary, full-time and part-time US employees who
work at least 20 hours per week, can be found here
[https://www.goldmansachs.com/careers/discover/2022-Benefits-Summary-US.pdf].
ABOUT GOLDMAN SACHS
At Goldman Sachs, we commit our people, capital and ideas to help our clients,
shareholders and the communities we serve to grow. Founded in 1869, we are a
leading global investment banking, securities and investment management firm.
Headquartered in New York, we maintain offices around the world.
We believe who you are makes you better at what you do. We're committed to
fostering and advancing diversity and inclusion in our own workplace and beyond
by ensuring every individual within our firm has a number of opportunities to
grow professionally and personally, from our training and development
opportunities and firmwide networks to benefits, wellness and personal finance
offerings and mindfulness programs. Learn more about our culture, benefits, and
people at GS.com/careers.
We’re committed to finding reasonable accommodations for candidates with special
needs or disabilities during our recruiting process. Learn more:
https://www.goldmansachs.com/careers/footer/disability-statement.html
© The Goldman Sachs Group, Inc., 2026. All rights reserved.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Senior Infrastructure Engineer - Data Protection
USAA · Tampa, Florida, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
Infrastructure Engineer I - Data Protection
USAA · Tampa, Florida, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Role information can change. Confirm current details on the original application page.
