Senior Site Reliability Engineer II
relx groupPlano, TX
today
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems AdministratorsIndustries
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesCustom Computer Programming ServicesAbout the role
Overview
In this role you design, build, and operate highly available systems in AWS, owning observability and automation to improve reliability. You will work closely with application teams to boost performance, deployability, and incident resilience. You’ll lead incident response and post-incident reviews, driving prevention and automation. The role offers hybrid or fully remote options and emphasizes ownership of production systems within a culture of automation and continuous improvement.
Compensation / Benefits Competitive salary and comprehensive benefits Flexible work location with hybrid or fully remote options Real ownership of production systems and reliability outcomesA culture that values automation, learning, and continuous improvement
Responsibilities Design, build, and operate highly available, scalable systems in AWSWrite, maintain, and review Terraform to provision and manage infrastructure Own and improve monitoring, alerting, and observability using Grafana, Pingdom, and Uptrends Participate in a rotating on-call schedule, responding to production incidents and driving issues to resolution Lead incident response, root cause analysis, and post-incident reviews with a focus on prevention and automation Define and manage SLOs, SLIs, and error budgets Build and improve CI/CD pipelines and operational workflows using Azure Dev Ops and Git Hub Work directly with application teams to improve reliability, performance, and deployability Automate manual operational tasks to reduce toil Maintain clear, actionable runbooks and documentation in Confluence Track work, incidents, and operational improvements using Jira and Service Now Mentor other engineers and help set SRE standards and best practices
Key requirements 5+ years of hands-on experience in SRE, Dev Ops, or Infrastructure Engineering roles Strong production experience in AWSSignificant hands-on experience with Terraform in real-world environments Experience operating monitoring and uptime platforms such as Grafana, Pingdom, and Uptrends Strong Linux systems, networking, and troubleshooting skills Experience supporting production systems through incident response and on-call rotations Proficiency with Git Hub and modern Git workflows Experience building or maintaining CI/CD pipelines with Azure Dev Ops Familiarity with ITSM and incident workflows using Service Now Strong written communication skills with experience documenting systems and processes in Confluence Ability to work independently in a remote or hybrid environmentstrong written communicationindependence in remote or hybrid environmentsmentoring others TerraformAWSGrafana
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Plano, TX
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
today
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE