About the role

Building and leading a new team, the full-time Manager of Site Reliability Engineering will oversee the availability, performance, and operability of the production control plane while managing incident response, monitoring, and configuration management in a remote environment.
Key Responsibilities: Build and lead the Production Site Reliability Engineering team, hiring and mentoring SREs and database engineers Own the availability and performance of the production web stack, leading incident response and driving system improvements Partner with engineering teams to ensure new services are production-ready and manage the re-architecture of the control plane
Required Qualifications: 10+ years of professional experience in site reliability engineering or infrastructure operations, including 2 years in a management role Deep experience operating production web stacks at scale and debugging performance issues Strong Linux systems knowledge and experience with configuration management at scale Proven track record of building monitoring and alerting systems and fostering a culture of observability Experience hiring and building engineering teams, with a focus on team growth and culture

Matching similar jobs

JOB OVERVIEW

Experience level

Manager

Location

Denver, CO

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE