About the role

Setting the reliability strategy for the platform, the full-time Principal Site Reliability Engineer will define deployment and operational standards for distributed systems, ensuring reliability and automation across customer environments while working remotely.
Key responsibilities: Own the reliability architecture of the platform, including deployment topology and automation for reproducible environments Define service level objectives and drive initiatives to meet them in collaboration with product and engineering leadership Lead major incidents and postmortems, ensuring corrective actions are implemented effectively
Required qualifications: 10+ years in infrastructure, SRE, or platform engineering with experience in large-scale distributed systems Expertise in designing and troubleshooting distributed systems and making reliability decisions under growth Deep experience with at least one major public cloud provider and proficiency in container orchestration Strong scripting and automation skills, along with familiarity in reading and debugging application code Experience operating within enterprise security and compliance frameworks

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Denver, CO

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE