About the role

To ensure operational readiness and reliability for production systems on a federal cloud platform, the full-time Cloud Site Reliability Engineer (SRE) will design automated remediation processes, define service-level objectives, and enforce infrastructure-as-code practices, working both onsite in Arlington, VA and remotely.
Key responsibilities: Design and implement automated systems for self-healing operations to enhance reliability Define data-driven service-level objectives and error budgets for critical services Enforce infrastructure-as-code standards and drive a containerization-first approach using Kubernetes
Required qualifications: Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience) 5+ years of SRE experience with demonstrated technical leadership Expertise in AWS and GovCloud, along with deep Kubernetes and Terraform knowledge at production scale Strong software engineering background in Python and/or Go Proven experience with observability platforms and incident management

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Denver, CO

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

yesterday

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE