Prepin
Log in
Patterson-UTI

engineering opportunity

Senior Site Reliability Engineer NEX

The Site Reliability Engineer will design, implement, and operate scalable and resilient systems on Google Cloud Platform while automating operational tasks to reduce toil. They will also partner with cross-functional teams to improve service reliability, manage incident response, and establish consistent engineering standards.

Houston, Texas, United StatesonsiteFULL_TIME

Posted

About the role

What will you do at Patterson-UTI?

Reliability Engineering

● Design, implement, and operate scalable, resilient, and highly available

systems on Google Cloud Platform.

● Improve service availability, latency, performance, scalability, and

operational resilience.

● Define, implement, and track service-level indicators, service-level

objectives, and error budgets.

● Perform capacity planning, performance analysis, and workload forecasting.

● Design and validate disaster recovery, backup, failover, and

service-restoration capabilities.

● Implement and maintain secure cloud networking, IAM, workload identities,

service accounts, and access-control practices.

● Partner with cybersecurity and identity teams to ensure infrastructure and

services follow organizational security standards.

● Monitor cloud consumption and optimize resource utilization, performance, and

cost efficiency.

● Identify operational risks and recommend improvements to cloud architecture

and service design.

Automation and Platform Engineering

● Build and maintain cloud infrastructure using Terraform or comparable

infrastructure-as-code tools.

● Automate repetitive operational activities and systematically identify,

measure, and reduce manual toil.

● Build reusable infrastructure modules, deployment patterns, and operational

tooling.

● Improve CI/CD pipelines to enable secure, repeatable, and reliable software

delivery.

Observability and Incident Management

● Develop actionable alerts that identify meaningful service degradation while

reducing alert fatigue and unnecessary operational noise.

● Create and maintain dashboards, runbooks, operational procedures, and

troubleshooting documentation.

● Participate in a sustainable on-call rotation supporting production systems.

● Respond to production incidents, coordinate service restoration, and lead

incident response when appropriate.

● Facilitate blameless postmortems and identify corrective and preventive

actions.

● Use incident and operational data to improve system design, automation,

monitoring, and response processes.

Collaboration and Service Ownership

● Partner with software engineering, data engineering, security, and product

teams to improve application reliability and production operations.

● Promote shared responsibility for production reliability between application

development and platform teams.

● Establish and document reliability standards, operational practices, and

reusable engineering patterns.

● Provide technical guidance and coaching on SRE, cloud, Kubernetes,

observability, and incident-management practices.

Required Knowledge, Skills, and Abilities

● Three or more years of experience in Site Reliability Engineering, platform

engineering, DevOps, cloud engineering, production software engineering, or a

similar role.

● Experience operating highly available systems in a 24/7 production

environment.

● Hands-on experience operating workloads on Google Cloud Platform or another

major public cloud platform.

● Strong experience managing compute, networking and data GCP services

workloads

● Strong experience with containerization and orchestration technologies,

including Docker and Kubernetes.

● Experience building and managing infrastructure with Terraform or a comparable

infrastructure-as-code tool.

● Proficiency in Python, Go, Java, or another comparable programming language.

● Experience implementing or operating CI/CD pipelines using GitHub

Actions,Azure DevOps, Bitbucket Pipelines, or comparable tools.

● Experience implementing observability using metrics, logs, traces, dashboards,

and alerts.

● Experience participating in on-call rotations, responding to incidents, and

contributing to postmortems.

● Understanding of SLIs, SLOs, error budgets, and other SRE principles.

● Ability to troubleshoot complex issues across application, infrastructure,

network, data, and cloud-service layers.

● Ability to communicate effectively with engineering teams, business

stakeholders, and operational personnel.

Minimum Qualifications

  • ● Bachelor’s degree in Computer Science, Information Technology, Engineering, or
  • a related field, or equivalent practical experience.
  • ● 3+ years of experience in Site Reliability Engineering, platform engineering,
  • cloud engineering, or DevOps.
  • ● 3+ years of experience operating production workloads in GCP.
  • ● Ability to understand and communicate in English at a level sufficient to
  • issue, receive, and respond to safety-related and operations-related
  • instructions.

Preferred Qualifications

  • ● Google Cloud and/or Kubernetes certifications.
  • ● Experience supporting data-intensive, streaming, analytics, or event-driven
  • platforms.
  • ● Experience establishing production-readiness, incident-management, or
  • reliability-review processes.
  • ● Experience working in the energy, oil and gas, industrial, IoT, field
  • operations, or other operationally critical industries.
  • ● Experience supporting technology environments that integrate cloud platforms
  • with remote sites, field equipment, industrial systems, or edge computing.
  • The Evolving Oil Field Demands Evolving Service Providers
  • NexTier is a leading provider of integrated completions that employs sustainable
  • practices and equipment to support our customers’ ESG goals while accelerating
  • production in the most demanding US land basins.
  • Patterson-UTI is committed to a workplace free from discrimination and
  • harassment, offering equal employment opportunities to all individuals
  • regardless of personal characteristics protected by law. Employees are
  • encouraged to report any concerns through multiple channels.

Which skills does this role require?

Site Reliability EngineeringGoogle Cloud PlatformTerraformDockerPythonGoJavaCI/CDObservabilityIncident ManagementInfrastructure as CodeCapacity PlanningAutomationCloud NetworkingService Level ObjectivesGitHub ActionsAzure DevOpsBitbucket PipelinesService Level IndicatorsError BudgetsPostmortemsIAMWorkload IdentitiesCybersecurityGCPAzure

Make your next move

Build a shortlist and prepare

Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.

Review the responsibilities and requirements before adding an opening to your shortlist.

Role information can change. Confirm current details on the original application page.

Product

AI Candidate AgentCompaniesBrowse JobsDeep ProfileSkill AssessmentOpportunity Matching
Prepin.ai

© 2026 Prepin | All rights reserved.