About the role

Find your next opportunity today. Principal Site Reliability Engineer Scottsdale, AZ85250Employment Type: Perm
Job Number: 25926 Remote Options: HybridJob Description
Principal Site Reliability Engineer
Position Overview: We are seeking an experienced Principal Site Reliability Engineer (SRE) to provide technical leadership across highly available, large-scale production environments. This role combines software engineering, systems engineering, cloud infrastructure, automation, DevOps, and production reliability to improve the resilience, scalability, performance, observability, and operational health of critical services. The Principal SRE will partner closely with Software Engineering, Platform Engineering, Cloud Infrastructure, DevOps, and other technology teams to ensure reliability and operational readiness are incorporated throughout the software development lifecycle. This is a senior individual contributor position with enterprise-level influence. The successful candidate will identify systemic reliability risks, establish technical direction, influence architecture and engineering practices, and help improve reliability capabilities across multiple engineering teams.
Key Responsibilities: Apply Site Reliability Engineering (SRE), software engineering, automation, and DevOps principles to improve how production services are built, tested, deployed, monitored, operated, and recovered. Establish and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, availability metrics, and service-health measurements. Design and enhance observability capabilities using metrics, logging, distributed tracing, monitoring, alerting, dashboards, and service-health instrumentation. Drive continuous improvement across CI/CD pipelines, Infrastructure as Code (IaC), cloud infrastructure, deployment practices, automation, testing, incident management, capacity planning, resilience, disaster recovery, and operational readiness. Analyze production environments to identify systemic reliability risks, performance bottlenecks, recurring incidents, and opportunities for automation. Translate production and operational experience into improvements in application code, architecture, infrastructure, tooling, automation, and engineering standards. Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle. Lead or participate in production incident response, troubleshooting, root cause analysis, service restoration, and blameless post-incident reviews. Provide technical leadership during critical production incidents and help improve incident response, escalation procedures, service restoration, and sustainable on-call practices. Reduce operational toil and manual intervention through software development, scripting, automation, reusable tooling, platforms, and engineering patterns. Apply data-driven analysis, experimentation, and engineering principles to validate assumptions and guide technical decisions. Establish and influence enterprise-level SRE, DevOps, cloud, reliability, and operational engineering standards and best practices. Mentor engineers and technical leaders while promoting knowledge sharing and sustainable engineering capabilities across the organization. Operate independently across complex, business-critical reliability and infrastructure challenges.
Required Qualifications:
  • 15+ years of relevant professional experience in one or more of the following areas: Site Reliability Engineering (SRE)
  • Software EngineeringSystems Engineering
  • Cloud Engineering
  • Platform Engineering
  • DevOps Engineering
  • Infrastructure EngineeringSystems Architecture
Strong experience with software development and/or scripting using one or more modern programming languages. Advanced understanding of software engineering principles, distributed systems, production environments, troubleshooting, automation, and observability. Experience designing, supporting, or improving highly available, scalable production systems and distributed applications. Experience with public cloud platforms and cloud-native architectures, preferably AWS.Strong knowledge of Linux/Unix systems, networking, infrastructure, application architecture, and production operations. Demonstrated experience diagnosing complex production issues and implementing sustainable technical solutions. Strong analytical, troubleshooting, problem-solving, communication, and cross-functional collaboration skills. Ability to provide technical direction and influence engineering practices across multiple teams and organizational boundaries.
Preferred Qualifications: Extensive hands-on experience with Amazon Web Services (AWS) or another major cloud platform, including Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).Experience with CI/CD pipelines and software delivery automation. Experience with Infrastructure as Code (IaC) technologies and practices. Experience with containers and container orchestration technologies. Strong experience with monitoring, logging, distributed tracing, dashboards, alerting, and observability platforms. Experience defining and managing SLIs, SLOs, error budgets, availability targets, and reliability metrics. Experience with incident management, root cause analysis, performance engineering, capacity planning, resilience testing, disaster recovery, and operational readiness. Experience creating reusable automation, tooling, platforms, frameworks, engineering patterns, or standards that improve engineering productivity and system reliability. Experience influencing architecture and technical strategy for large-scale or business-critical production systems. Demonstrated ability to mentor senior engineers and improve technical capabilities across engineering organizations. Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical discipline, or equivalent practical experience.
Key Technical Skills / ATS Keywords: Site Reliability Engineering (SRE), AWS, Cloud Computing, DevOps, Software Engineering, Systems Engineering, Platform Engineering, Distributed Systems, Production Reliability, High Availability, Scalability, Resilience, Observability, Infrastructure as Code (IaC), CI/CD, Automation, Linux, Unix, Networking, Containers, Container Orchestration, Monitoring, Logging, Distributed Tracing, Alerting, SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis, Production Support, Performance Engineering, Capacity Planning, Disaster Recovery, Resilience Testing, Operational Readiness, Cloud Architecture, Production Operations, Software Development, Scripting, Troubleshooting, Technical Leadership. Share This Job: to save this search and get notified of similar positions. Related Jobs: There are currently no related jobs. Please sign up for Job Alerts. Loading... to save this search and get notified of similar positions. About Scottsdale, AZReady to embark on a career adventure in Scottsdale, Arizona? This vibrant city nestled in the heart of Maricopa County offers a perfect blend of desert landscapes, world-class resorts, and a thriving arts scene, promising endless growth opportunities for job seekers. From exploring the iconic Camelback Mountain to indulging in Southwestern cuisine at local favorites like The Mission, Scottsdale is a place where career aspirations and quality of life harmoniously coexist. With nearby attractions like the Scottsdale Museum of Contemporary Art, Taliesin West, and the Scottsdale Stadium - home to the San Francisco Giants during spring training - there is no shortage of unique experiences waiting to be discovered. Dive into our job listings today and unlock your full potential in this dynamic and enchanting city! Are you sure you want to apply for this job? Please take a moment to verify your personal information and resume are up-to-date before you apply. Snooze message for 30 days By checking this box, you will snooze this confirmation message 30 days, and your application will be automatically submitted with your saved information. If you wish to edit your information please visit My Profile Section

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Scottsdale, AZ

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

2 days ago

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE