About the role
What will you do at Marriott International Hotels, Inc.?
OB SUMMARY
The Systems Engineer - Site Reliability Engineering (SRE) is responsible for the
reliability, scalability, and performance of mission-critical cloud and on-prem
services that support millions of Marriot customers globally. This role involves
overseeing incident management, driving automation efforts, and working closely
with cross-functional teams to ensure alignment between SRE strategy and
business objectives. Partners closely with Product Teams, Applications teams,
Infrastructure, and the broader Applications and Infrastructure Delivery teams
to develop key metrics and KPIs to improve applications stability, availability
and performance. The ideal candidate will bring strong communication skills,
collaborating with key stakeholders across the company to optimize cloud
infrastructure and uphold the highest standards of operational excellence in a
dynamic, fast-paced environment
CANDIDATE PROFILE
Required Education and Experience
* Undergraduate degree in an engineering or computer science discipline and/or
equivalent experience/certification
* 5+ years of experience as a Site Reliability Engineer (SRE), building and
managing highly available and mission critical systems
* Expertise in AWS services including designing highly available, multi-AZ and
multi-region architectures including:
* Compute: EC2, Auto Scaling, Lambda
* Containers: EKS (Mandatory), ECS (good to have)
* Networking: VPC, subnets, route tables, NAT gateways, Transit Gateway
* Security: IAM roles/Policies, KMS, Secret manager
* Storage and Databases: S3, EBS, EFS, RDS, DocumentDB.
* Deep understanding of SRE practices such as Service Level Objectives, Error
Budgets, Toil Management, Observability & Monitoring, Blameless Postmortems,
Incident Response Process, Capacity Planning
* Strong understanding of Cloud Security best practices and responsibility
model.
* Experience driving cloud cost optimization initiatives (rightsizing, reserved
instances, autoscaling strategies, cost observability)
* Proven automation and programming experience in one or more of the following
languages: Python, Bash, PowerShell
* Strong working knowledge of modern, continuous development techniques and
pipelines (Agile, Kanban, Jira, CI/CD, Helm, Harness, Jenkins, Git,
Artifactory, Vault)
* Production level expertise with containerization orchestration engines such
as Kubernetes (EKS, AKS, ACK)
* Hands-on experience with service mesh technologies to enable secure and
resilient service communication, including mTLS, traffic shaping, and policy
enforcement.
* Strong experience troubleshooting API-related issues in distributed systems,
including latency, authentication/authorization failures, rate limiting, and
upstream/downstream dependency failures.
* Ability to analyze API traffic and debug issues using logs, traces, and
metrics.
* Experience with Infrastructure as Code (Iac) tools like Terraform and
CloudFormation.
* Experience with configuration management and automation tools such as
Ansible.
* Deep expertise and hands-on experience with Linux administration (RHEL,
Ubuntu, CentOS, AWS Linux)
* Solid understanding of Virtualization Technologies (VMware vSphere, KVM etc)
* Strong understanding of networking fundamentals such as Load Balancing,
Firewalls, Security Groups, NACLs, TCP/IP, DNS, HTTP/HTTPS, SSL/TLS etc
* Deep understanding and/or experience with Cloud Native, Relational and NoSQL
databases like RDS, MySQL, PostgreSQL, Cassandra or Couchbase
* Strong experience designing and implementing end-to-end observability
solutions across metrics, logs, and traces using tools like Prometheus,
Grafana, ELK Stack, and OpenTelemetry.
* Proven ability to define SLIs/SLOs, build actionable alerting systems, and
leverage telemetry data for incident response, root cause analysis, and
performance optimization.
* Experience with deploying, monitoring, and troubleshooting large-scale,
distributed applications in cloud environments such as AWS
* Experience in vulnerability management, OS hardening, patching, security
compliance of infrastructure, applications and databases
* Experience in implementing OS and cloud hardening guidelines and perform
regular vulnerability remediation.
* Familiarity with security frameworks such as ISO27001, SOCII, PCI-DSS, and/or
HIPAA
* 5+ years progressive technology experience
* 3+ years’ experience in the operational support of critical solutions in
large scale environments and organizations with specific experience in
* Windows Servers
* RHEL
* VMWare (vCenter and ESXi Hypervisor)
* Ability to work outside of normal business hours and days in a 24x7x365
environment.
* Some travel required
Preferred:
* CI/CD pipeline technologies such as Git, Docker Trusted Registry,
Artifactory, Hashicorp Vault, Maven, etc.
* Security Protocols like SSL, SAML, LDAP etc.
* AIX Administration
* Automation and Scripting languages (Ansible, PowerShell, VB, ShellScript)
* Strong organizational, written and verbal communication skills
* ITIL v4 certification
* Windows and Redhat Certification
* Experience delivering technology solutions in a fast-paced, deadline driven
enterprise environment
* Experience learning and applying new technologies to solve business needs
* Excellent understanding of change management, testing requirements,
techniques, and tools to ensure high availability of systems
* Experience in researching emerging technologies and trends, standards, and
products
CORE WORK ACTIVITIES
* Provide support for priority incidents as directed by SRE leader
* Collaborates through the incident with key team members (network,
application, etc.) to engage service providers and other stakeholders to
identify problem root cause and drive service restoration
* Provides closure to incidents or initiates a process-based, hand-off to next
shift
* Works with support vendors and providers to ensure proper global coverage,
phone support, parts replacement and smart hands
* Work to create and mature incident response processes
* Provides oversight to service providers or less experienced engineers
* Drive the objectives associated with Problem Management; such as customer
communication and Root Cause Analysis reports
* Trains and/or mentors other team members, and peers as appropriate
* Identifies opportunities to enhance the service delivery, operations and
continual service improvement processes
* Identifies solutions that may contribute to greater stability and reliability
of property infrastructure
* Develop implementation plans, test plans, and timelines for projects and
tasks
Delivering Technology
* Create and enhance administrative, operational and technical policies and
procedures, adopting best practice guidelines, standards and procedures for
employees, contractors and vendor engagements
* Maintains a proper balance between business and operational risk
* Establishes cadence of communication with other organizations to stay abreast
of deployment and production support activities
* Works in a concerted effort with application development and engineering
teams to resolve complex issues
* Provides oversight, collaboration, provisioning, management and maintenance
of technology products and service alternatives that improve the production
services environment
* Responsible for the establishment and continuous development of monitoring
and alerting for all production environments
* Contributes to continuous improvement of internal processes
* Attends training to ensure skillset and tools support the production
environments and deliver on project commitments
* Performs quantitative and qualitative analyses for operational availability
to promote a zero-defect environment
* Facilitates achievement of expected deliverables and obligations of Services
Providers
* Assists operational teams in system updates & upgrades
* Provides consultation for routine systems development
* Ensures early warning to the business stakeholder executives regarding
degraded or missed service levels
Service Provider Management
* Actively coordinates with IT service providers and vendors to bring incidents
to resolution
* Manages Service Providers with a focus on continuous service improvement and
service restoration
* Monitors, manages and leads Service Provider outcomes required to ensure
operational availability and a zero-defect production environment
* Consults with internal Service Management & external Service Providers on
performance, business reporting, analytics metrics and business value
dashboards
Maintaining Goals
* Submits reports in a timely manner, ensuring delivery deadlines are met.
* Promotes the documenting of project progress accurately.
* Provides input and assistance to other teams regarding projects.
Managing Work, Projects, and Policies
* Manages and implements work and projects as assigned.
* Generates and provides accurate and timely results in the form of reports,
presentations, etc.
* Analyzes information and evaluates results to choose the best solution and
solve problems.
* Provides timely, accurate, and detailed status reports as requested.
Demonstrating and Applying Discipline Knowledge
* Provides technical expertise and support to people inside and outside of the
department.
* Demonstrates knowledge of job-relevant issues, products, systems, and
processes.
* Demonstrates knowledge of function-specific procedures.
* Keeps up-to-date technically and applies new knowledge to job.
* Uses computers and computer systems (including hardware and software) to
enter data and/ or process information.
Delivering on the Needs of Key Stakeholders
* Understands and meets the needs of key stakeholders.
* Develops specific goals and plans to prioritize, organize, and accomplish
work.
* Determines priorities, schedules, plans and necessary resources to ensure
completion of any projects on schedule.
* Collaborates with internal partners and stakeholders to support
business/initiative strategies
* Communicates concepts in a clear and persuasive manner that is easy to
understand.
* Generates and provides accurate and prompt results in the form of reports,
presentations, etc.
* Demonstrates an understanding of business priorities
At Marriott International, we are dedicated to being an equal opportunity
employer, welcoming all and providing access to opportunity. We actively foster
an environment where the unique backgrounds of our associates are valued and
celebrated. Our greatest strength lies in the rich blend of culture, talent, and
experiences of our associates. We are committed to non-discrimination on any
protected basis, including disability, veteran status, or other basis protected
by applicable law.
All positions offer a 401(k) plan, stock purchase plan, discounts at Marriott
properties, commuter benefits, employee assistance plan, and childcare
discounts.
Benefits
are subject to terms and conditions, which may include
rules regarding eligibility, enrollment, waiting period, contribution, benefit
limits, election changes, benefit exclusions, and others. Click here
[https://life.marriott.com/wp-content/uploads/2025/09/benefitsoverviewp_2025edits_8.19.25.pdf] to
learn more.
Full-time positions also offer coverage for medical, dental, vision, health care
flexible spending account, dependent care flexible spending account, life
insurance, disability insurance, accident insurance, adoption expense
reimbursements, paid parental leave and educational assistance.
Washington Applicants Only: Employees will accrue paid sick leave, 0.077 PTO
balance for every hour worked and be eligible to receive a minimum of 9 holidays
annually.
Marriott Headquarters (HQ) supports a hybrid work environment that helps
associates Be Connected. HQ based roles are generally hybrid and require regular
in-office presence at the assigned location, including for candidates who live
within commuting distance of HQ for a role listed as remote. Remote roles will
be clearly identified in the job posting.
Marriott International is the world’s largest hotel company, with more brands,
more hotels and more opportunities for associates to grow and succeed. Be where
you can do your best work, begin your purpose, belong to an amazing global
team, and become the best version of you.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Senior Infrastructure Engineer - Data Protection
USAA · Tampa, Florida, United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Infrastructure Engineer I - Data Protection
USAA · Tampa, Florida, United States
Role information can change. Confirm current details on the original application page.
