Senior Site Reliability Engineer

castleton commodities internationalThe Woodlands, TX

today

Occupations

Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems Administrators

Industries

Computer Systems Design ServicesOther Computer Related ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
APPLY NOW

About the role

Overview As Senior Site Reliability Engineer, you will drive reliability, scalability, and operational excellence for critical infrastructure. You’ll partner with Engineering, Security, and Infrastructure teams to design cloud-native, resilient architectures and implement IaC and CI/CD standards. You define recovery objectives (RTO/RPO), shape BCP/DR plans, and lead structured testing to prove readiness. You’ll own observability, incident response, and capacity planning to reduce MTTR and enable safe, automated deployments. Compensation / Benefits Competitive comprehensive medical, dental, retirement and life insurance Employee assistance & wellness programs Parental and family leave policies Tuition assistance & reimbursement Quarterly Innovation & Collaboration Awards Competitive paid time off Responsibilities Own and improve service reliability via SLO/SLI, error budgets, and best practices Design, implement, and maintain observability (monitoring, logging, tracing, alerting)Lead incident response including on-call improvements, runbooks, and post-incident reviews Collaborate with application teams to improve performance, capacity planning, and resiliency Design and operate highly available cloud architectures (multi-AZ, multi-region)Implement resilient patterns across compute, storage, networking, and managed services Develop, standardize, and optimize IaC modules (Terraform, Cloud Formation, CDK)Build and optimize CI/CD pipelines for automated deployments Ensure environment consistency across dev/test/stage/prod and drift remediation Collaborate on BCP/DR strategies, testing, and recovery documentation Produce metrics and reporting on DR readiness and continuous improvement actions Key requirements 7+ years of experience in SRE, Dev Ops, Platform Engineering, or Systems Engineering in production Strong proficiency with observability platforms (Datadog, Prometheus/Grafana, ELK/Open Search, Nagios, Nimsoft)Strong hands-on AWS experience Experience with IaC (Terraform and/or Cloud Formation/CDK)Strong CI/CD and automation background Experience defining and validating RTO/RPO and implementing BCP/DR with testing Kubernetes experience with auto-scaling (EKS, ECS, or on-prem)Strong Linux fundamentals and networking knowledge Scripting/programming in Python, Go, Bash, or similar Ability to write clear operational docs, runbooks, and post-incident reports Ability to work in fast-paced, high-intensity environments and outside hours when required Collaborative mindset Strong communication and documentation Problem-solving with structured, data-driven approaches Datadog Prometheus/GrafanaELK/Open Search

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

The Woodlands, TX

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE