Staff Site Reliability Engineer
virtual vocationsNew York, NY
today
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems AdministratorsIndustries
Computer Systems Design ServicesCustom Computer Programming ServicesSoftware PublishersAbout the role
As the first dedicated Staff Site Reliability Engineer in a fully remote capacity, the successful candidate will manage reliability metrics, enhance incident response processes, and foster a reliability-focused mindset across engineering teams, ensuring operational excellence and measurable reliability standards.
Key responsibilities Define and implement SLIs and SLOs for critical request paths, ensuring visibility and accountability among teams
Strengthen the incident lifecycle by improving detection, response, and postmortem processes while driving alert quality and escalation design
Coach teams on reliability practices, embedding a culture of operational excellence and deliberate failure testing Required qualifications 10+ years of engineering experience, including 3+ years as a Site Reliability Engineer or in a similar reliability-focused role
Proven expertise in SLI/SLO design and error budgets, with a track record of team adoption
Strong incident leadership experience, having managed high-severity incidents and improved organizational learning
Hands-on experience with distributed systems and proficiency in Kubernetes, AWS, and modern observability tools
Ability to read and write production code (Go, Type Script, or similar) and familiarity with infrastructure as code
Matching similar jobs
JOB OVERVIEW
Experience level
Manager
Location
New York, NY
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
today
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE