Senior Site Reliability Engineer NEX
pattersonutiThe Woodlands, TX
yesterday
Occupations
Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsSoftware DevelopersIndustries
Computer Systems Design ServicesOther Computer Related ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesAbout the role
Overview
In this role you will design, operate, and scale reliable cloud-based systems on Google Cloud Platform, focusing on availability, performance, and resilience. You will define SLOs, error budgets, and DR strategies, while collaborating with security and platform teams to uphold standards. You’ll automate infrastructure, improve CI/CD, and drive observability and incident response. This is an opportunity to shape production reliability for critical energy-industry workloads at scale.
Responsibilities Design, implement, and operate scalable, highly available systems on Google Cloud Platform Define and monitor SLOs, error budgets, and service indicators Perform capacity planning, performance analysis, and workload forecasting Design disaster recovery, backup, failover, and service restoration capabilities Manage secure cloud networking, IAM, service accounts, and access controls Collaborate with cybersecurity and identity teams to enforce security standards Improve cloud resource utilization, performance, and cost efficiency Identify operational risks and propose cloud architecture improvements Build and maintain infrastructure with Terraform or similar IaC tools Automate repetitive tasks and reduce manual toil Enhance CI/CD pipelines for secure, repeatable software delivery Develop actionable alerts and dashboards to reduce noise Lead incident response, blameless postmortems, and preventive actions Provide guidance on SRE, cloud, Kubernetes, observability, and incident management Promote shared responsibility for production reliability across teams
Key requirements Three or more years in SRE/platform engineering/Dev Ops/related roles Experience operating highly available 24/7 production systems Hands-on experience with Google Cloud Platform or another major cloud Strong compute, networking, and data GCP workloads experience Kubernetes and Docker proficiency Terraform or comparable IaC tooling experience Proficiency in Python, Go, Java, or equivalent language Experience with CI/CD pipelines (Git Hub Actions, Azure Dev Ops, Bitbucket Pipelines, or similar)Observability via metrics, logs, traces, dashboards, and alerts Experience in on-call rotations and postmortems Understanding of SLIs, SLOs, and error budgets Ability to troubleshoot across app, infra, network, data, and cloud services Effective communication with engineers, stakeholders, and operators Bachelor's degree or equivalent practical experience English proficiency for safety/operational instructionscollaborationclear communicationincident management Google Cloud Platform Kubernetes Docker
Matching similar jobs
JOB OVERVIEW
Experience level
Manager
Location
The Woodlands, TX
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
yesterday
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE