About the role

GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation, reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and is a hybrid remote/onsite position.
ResponsibilitiesKey Responsibilities: Infrastructure & AutomationDesign, deploy, and manage cloud infrastructure using Infrastructure as Code (IaC) principlesDevelop and maintain Terraform modules for AWS and Azure environmentsCreate and manage Ansible playbooks for configuration management and application deploymentImplement CI/CD pipelines using GitHub Actions to automate build, test, and deployment processesImplement GitOps workflows for declarative infrastructure and application deliveryBuild self-service tools and platforms to enable development teamsReliability & PerformanceEstablish and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) Implement comprehensive monitoring, logging, and alerting solutionsConduct capacity planning and performance tuningPerform root cause analysis and implement preventive measuresDesign and execute chaos engineering experiments to validate system resilienceDisaster Recovery & Business ContinuityDesign and implement disaster recovery strategies across multi-cloud environmentsDevelop and maintain backup and restore proceduresCreate and test business continuity plansImplement automated failover mechanismsDocument recovery time objectives (RTO) and recovery point objectives (RPO) Cloud OperationsManage and optimize AWS services (EC2, S3, RDS, Lambda, ECS, EKS, CloudWatch, etc.)Manage and optimize Azure services (VMs, Storage, SQL Database, AKS, Monitor, etc.)Implement cost optimization strategies and resource taggingEnsure security best practices and compliance requirementsManage identity and access management (IAM) policiesCollaboration & LeadershipParticipate in on-call rotation and incident responseCollaborate with development teams on architecture and design decisionsMentor team members on SRE practices and toolsDocument systems, processes, and runbooksDrive continuous improvement initiativesQualificationsRequiredEducation and ExperienceBachelor's Degree with 12+ yrs experienceTechnicalSkillsCloud Platforms: 3+ years of hands-on experience with AWS and AzureInfrastructure as Code: Expert-level proficiency with TerraformConfiguration Management: Strong experience with AnsibleScripting: Proficiency in Python, Bash, or PowerShellContainerization: Experience with Docker and KubernetesVersion Control: Strong Git and GitHub workflow knowledgeGitOps: Experience implementing GitOps practices and workflowsMonitoring Tools: Experience with Prometheus, Grafana, ELK Stack, or similarCI/CD: Hands-on experience with GitHub Actions, Jenkins, GitLab CI, or Azure DevOpsCore CompetenciesDeep understanding of Microsoft/Linux systems administrationStrong networking knowledge (TCP/IP, DNS, load balancing, VPN) Experience with database administration (PostgreSQL, MySQL, SQL Server) Knowledge of security best practices and compliance frameworksUnderstanding of microservices architecture and distributed systemsExperience with disaster recovery planning and executionSoftSkillsExcellent problem-solving and analytical abilitiesStrong communication skills, both written and verbalAbility to work independently and in team environmentsCustomer-focused mindset with emphasis on reliabilityAdaptability to rapidly changing technologies and requirementsClearance Required: Active Secret with the ability to obtain and hold DEA suitabilityPreferredQualificationsAWS Certified Solutions Architect or SysOps AdministratorAzure Administrator or Solutions Architect certificationCertified Kubernetes Administrator (CKA) HashiCorp Certified: Terraform AssociateGitHub Certified or demonstrated expertise with GitHub EnterpriseExperience with service mesh technologies (Istio, Linkerd) Knowledge of observability platforms (Datadog, New Relic, Dynatrace) Experience with GitOps tools and practices (ArgoCD, Flux, GitHub Actions for GitOps) Familiarity with compliance frameworks (SOC 2, HIPAA, FedRAMP) Previous experience in a DevOps or Platform Engineering rolePosted Salary RangeUSD $210,000.00 - USD $230,000.00 /Yr.

Matching similar jobs

JOB OVERVIEW

Salary

$230,000.00 /Yr

Experience level

Lead

Location

Arlington, VA

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

17 days ago

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE