About the role

About The Role: The Site Reliability Engineer owns the uptime, performance, and scalability of production systems handling thousands of requests per second. The role blends software engineering with systems thinking - eliminating toil through automation, designing infrastructure that fails gracefully, and running incident response that keeps customers informed and services recovering fast. This is a hands‑on role on a platform team that other engineering teams depend on. You will shape how services are deployed, monitored, and scaled across Kubernetes and cloud infrastructure, and your work directly determines whether engineers can ship safely and customers stay online.
Key Responsibilities: Build and maintain scalable infrastructure on AWS/GCP using Terraform, treating everything as code with peer‑reviewed modules and CI‑driven provisioning Own SLOs and error budgets for critical services - define SLIs, build dashboards in Grafana/Datadog, and drive reliability decisions across product teams Design and operate Kubernetes clusters: autoscaling, resource tuning, pod disruption budgets, and Helm‑based service deployment standards Lead incident response as an on‑call escalation point - run incident command, write blameless postmortems, and drive remediation work to closure Automate toil away with Python, Bash, and Go - self‑healing runbooks, capacity management, and cost optimization tooling Harden production systems: implement network policies, secrets management, least‑privilege IAM, and vulnerability remediation pipelines Partner with development teams on production readiness - review designs for reliability, set deployment standards, and improve observability coverage What We Are Looking For3–7 years of experience in SRE, DevOps, or infrastructure engineering, including operating large‑scale production systems with real on‑call responsibilities Deep hands‑on experience with Kubernetes in production - operating clusters, debugging workloads, and tuning for performance and cost Strong infrastructure‑as‑code skills with Terraform (or similar) and experience with CI/CD pipelines (GitHub Actions, GitLab CI, or ArgoCD) Solid grasp of distributed systems fundamentals: load balancing, service meshes, queues, caching strategies, and failure modes Production experience with observability stacks: Prometheus, Grafana, Datadog, or equivalent - including building actionable alerts, not noisy ones BS in Computer Science or equivalent practical experience; scripting proficiency in Python, Go, or Bash
Bonus: experience with chaos engineering, multi‑region architectures, service mesh (Istio/Linkerd), or database reliability (PostgreSQL, MySQL)

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Phoenix, AZ

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE