Senior Site Reliability Engineer NEX

pattersonutiThe Woodlands, TX

yesterday

Occupations

Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsSoftware Developers

Industries

Computer Systems Design ServicesOther Computer Related ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
APPLY NOW

About the role

Overview In this role you will design, operate, and scale reliable cloud-based systems on Google Cloud Platform, focusing on availability, performance, and resilience. You will define SLOs, error budgets, and DR strategies, while collaborating with security and platform teams to uphold standards. You’ll automate infrastructure, improve CI/CD, and drive observability and incident response. This is an opportunity to shape production reliability for critical energy-industry workloads at scale. Responsibilities Design, implement, and operate scalable, highly available systems on Google Cloud Platform Define and monitor SLOs, error budgets, and service indicators Perform capacity planning, performance analysis, and workload forecasting Design disaster recovery, backup, failover, and service restoration capabilities Manage secure cloud networking, IAM, service accounts, and access controls Collaborate with cybersecurity and identity teams to enforce security standards Improve cloud resource utilization, performance, and cost efficiency Identify operational risks and propose cloud architecture improvements Build and maintain infrastructure with Terraform or similar IaC tools Automate repetitive tasks and reduce manual toil Enhance CI/CD pipelines for secure, repeatable software delivery Develop actionable alerts and dashboards to reduce noise Lead incident response, blameless postmortems, and preventive actions Provide guidance on SRE, cloud, Kubernetes, observability, and incident management Promote shared responsibility for production reliability across teams Key requirements Three or more years in SRE/platform engineering/Dev Ops/related roles Experience operating highly available 24/7 production systems Hands-on experience with Google Cloud Platform or another major cloud Strong compute, networking, and data GCP workloads experience Kubernetes and Docker proficiency Terraform or comparable IaC tooling experience Proficiency in Python, Go, Java, or equivalent language Experience with CI/CD pipelines (Git Hub Actions, Azure Dev Ops, Bitbucket Pipelines, or similar)Observability via metrics, logs, traces, dashboards, and alerts Experience in on-call rotations and postmortems Understanding of SLIs, SLOs, and error budgets Ability to troubleshoot across app, infra, network, data, and cloud services Effective communication with engineers, stakeholders, and operators Bachelor's degree or equivalent practical experience English proficiency for safety/operational instructionscollaborationclear communicationincident management Google Cloud Platform Kubernetes Docker

Matching similar jobs

JOB OVERVIEW

Experience level

Manager

Location

The Woodlands, TX

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

yesterday

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE
Senior Site Reliability Engineer NEX at pattersonuti | Johnson Jobs