Sr. Site Reliability Engineer, AI Infrastructure (Starshield)
spacexAlexandria, VA
today
Occupations
Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsSoftware DevelopersIndustries
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesOther Computer Related ServicesAbout the role
Overview
In this role you will design, operate, and scale Starshield’s GPU-centric infrastructure to support national security missions. You’ll collaborate with AI engineers and cross‑functional teams to deliver highly available, on‑premise compute services and AI clusters at scale. You drive automation, reliability, and maintainability of core systems from design through deployment and operation. This position offers a chance to shape mission-critical infrastructure for a globally deployed satellite constellation.
Compensation / Benefitsstock or long-term incentivesdiscretionary bonusesmedical, vision, dental coverage 401(k) planpaid parental leavepaid vacation and holidays
Responsibilities Manage GPU/CPU infrastructure deployments in Top Secret data centers Provide GPU-as-a-service for external customers on bare metal and virtualized platforms Design and productize AI clusters at large scale (100k+ GPUs)Develop automation for on‑premise Kubernetes/AI clusters and OSsDeploy and manage core infrastructure (databases, monitoring, distributed storage)Collaborate with AI engineers to create scalable, operable products Oversee full lifecycle of services from design to refinement Maintain monitoring and alerting for high availability Identify improvements and implement high-availability solutions Mentor and train junior engineers Lead the team to technical excellence as a senior engineer
Key requirements Bachelor’s degree in CS/IT/engineering or 5+ years with Linux OR 7+ years in software/Dev Ops/SE5+ years of experience with Kubernetes 5+ years of experience managing Linux operating systems Experience with Terraform, Ansible, or other infrastructure tools Experience with containerization (OCI/Kubernetes)Scripting in Bash, Python, or similar languages Development experience in Python, C++, or GoExcellent communications skills with customers, peers, and management Mentorship and willingness to train junior engineers Strong problem-solving and collaboration abilities Kubernetes cluster management Terraform and Ansible Linux system administration
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Alexandria, VA
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
today
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE