Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

spacexAlexandria, VA

today

Occupations

Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsSoftware Developers

Industries

Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesOther Computer Related Services
APPLY NOW

About the role

Overview In this role you will design, operate, and scale Starshield’s GPU-centric infrastructure to support national security missions. You’ll collaborate with AI engineers and cross‑functional teams to deliver highly available, on‑premise compute services and AI clusters at scale. You drive automation, reliability, and maintainability of core systems from design through deployment and operation. This position offers a chance to shape mission-critical infrastructure for a globally deployed satellite constellation. Compensation / Benefitsstock or long-term incentivesdiscretionary bonusesmedical, vision, dental coverage 401(k) planpaid parental leavepaid vacation and holidays Responsibilities Manage GPU/CPU infrastructure deployments in Top Secret data centers Provide GPU-as-a-service for external customers on bare metal and virtualized platforms Design and productize AI clusters at large scale (100k+ GPUs)Develop automation for on‑premise Kubernetes/AI clusters and OSsDeploy and manage core infrastructure (databases, monitoring, distributed storage)Collaborate with AI engineers to create scalable, operable products Oversee full lifecycle of services from design to refinement Maintain monitoring and alerting for high availability Identify improvements and implement high-availability solutions Mentor and train junior engineers Lead the team to technical excellence as a senior engineer Key requirements Bachelor’s degree in CS/IT/engineering or 5+ years with Linux OR 7+ years in software/Dev Ops/SE5+ years of experience with Kubernetes 5+ years of experience managing Linux operating systems Experience with Terraform, Ansible, or other infrastructure tools Experience with containerization (OCI/Kubernetes)Scripting in Bash, Python, or similar languages Development experience in Python, C++, or GoExcellent communications skills with customers, peers, and management Mentorship and willingness to train junior engineers Strong problem-solving and collaboration abilities Kubernetes cluster management Terraform and Ansible Linux system administration

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Alexandria, VA

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE