Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY

ardan labsDenver, CO

today

Occupations

Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems Administrators

Industries

Computer Systems Design ServicesCustom Computer Programming ServicesSoftware Publishers
APPLY NOW

About the role

We are looking for a Senior Site Reliability Engineer to join a platform team and take ownership of both the reliability of an internal developer platform and the experience of the application teams building on it. This is a hands-on senior individual contributor role for an engineer who enjoys solving complex infrastructure problems, operating Kubernetes at scale, and working directly with teams to diagnose and resolve issues. You will work across Amazon EKS and Red Hat Open Shift on AWS, a curated Helm chart catalog, Git Ops pipelines, Terraform, and a multi-account AWS environment built on Control Tower. Our workloads operate under HIPAA requirements, making security hygiene, patch currency, and operational reliability ongoing engineering priorities. Approximately 60% of your time will focus on reliability engineering and 40% on forward-deployed work with application, security, network, and other technical teams.
What You'll Do: Own the health and reliability of Kubernetes management and workload clusters across sandbox, development, QA, and production environments. Diagnose and permanently resolve Git Ops delivery and reconciliation failures. Manage version and component adoption to ensure fixes and improvements are consistently deployed across clusters. Develop and maintain Terraform and infrastructure pipelines across a multi-account AWS environment. Work with cloud identity, resource policies, and key management to diagnose and resolve complex permission issues. Partner directly with application teams to troubleshoot manifests, managed resource claims, secrets, certificates, ingress, and other platform issues. Build documentation, guides, and guardrails that help application teams become increasingly self-sufficient. Own and improve the telemetry path into the observability platform. Partner with security, networking, vendors, and other technical teams when the platform is involved in an incident. Participate in incident response and drive problems through to durable resolution. Lead technical initiatives across teams that do not report to you.
What We're Looking For: 8+ years of experience in infrastructure, platform engineering, or site reliability engineering. At least 3 years of production Kubernetes experience supporting teams beyond your own. Strong hands-on experience with production Git Ops, including diagnosing conflicts between declared and live state. Experience using Terraform across multiple cloud accounts or environments. Deep understanding of cloud identity and resource policies, including troubleshooting permissions that appear correct but are still denied. Demonstrated ability to lead technical work and influence teams without direct authority. Strong troubleshooting, systems thinking, and communication skills. Comfortable working directly with application and engineering teams rather than operating solely behind the scenes.
Nice to Have: The following are not required, but would strengthen your application: Red Hat Open Shift / ROSA experience Experience working in regulated environments Production experience with Istio or another service mesh Forward-deployed, field engineering, implementation, solutions, or customer-facing engineering experience Ownership of Open Telemetry, Prometheus, or commercial APM platforms Go or another compiled language used for infrastructure/tooling Published technical writing, such as incident reports, design documents, or technical/customer-facing documentation Why This Role Is Different This isn't a role where success is measured by how many tools you know. The platform is highly automated and largely self-healing. The difficult problems are often about figuring out where the problem actually belongs, determining the right durable fix, and building agreement between teams that don't report to one another. We're looking for an engineer who can operate at both levels: deep technical infrastructure expertise and strong cross-team engagement.

Matching similar jobs

JOB OVERVIEW

Experience level

Manager

Location

Denver, CO

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE
Senior Site Reliability Engineer, Forward Deployed - Remote USA...