Staff Site Reliability Engineer
virtual vocationsDenver, CO
Staff Site Reliability Engineer
virtual vocationsDenver, CO
today
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems AdministratorsAbout the role
As the first dedicated Staff Site Reliability Engineer in a fully remote capacity, the successful candidate will manage reliability metrics, enhance incident response processes, and foster a reliability-focused mindset across engineering teams, ensuring operational excellence and measurable reliability standards.
Key responsibilities:
Define and implement SLIs and SLOs for critical request paths, ensuring visibility and accountability among teams
Strengthen the incident lifecycle by improving detection, response, and postmortem processes while driving alert quality and escalation design
Coach teams on reliability practices, embedding a culture of operational excellence and deliberate failure testing
Required qualifications:
10+ years of engineering experience, including 3+ years as a Site Reliability Engineer or in a similar reliability-focused role
Proven expertise in SLI/SLO design and error budgets, with a track record of team adoption
Strong incident leadership experience, having managed high-severity incidents and improved organizational learning
Hands-on experience with distributed systems and proficiency in Kubernetes, AWS, and modern observability tools
Ability to read and write production code (Go, TypeScript, or similar) and familiarity with infrastructure as code
Matching similar jobs
JOB OVERVIEW
Experience level
Manager
Location
Denver, CO
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
today
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE