Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
motion recruitmentMt Laurel, NJ
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
motion recruitmentMt Laurel, NJ
8 days ago
Occupations
Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsSoftware DevelopersIndustries
Computer Systems Design ServicesOther Computer Related ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesAbout the role
Mount Laurel, New Jersey 100% Remote Full Time$150k - $200kRemote (USA) | Full-Time | Site Reliability Engineer Join a rapidly growing B2B AI infrastructure company powering large-scale machine learning and AI workloads for more than one million developers worldwide. As a Site Reliability Engineer, you'll help improve the reliability, scalability, and performance of a cloud platform built on Linux, Kubernetes, distributed systems, GPU infrastructure, observability, and automation technologies. This full-time remote opportunity offers the chance to work on critical infrastructure supporting AI applications on a global scale. As the company continues to scale its AI infrastructure platform, reliability has become a critical business function. This role sits at the center of that effort, partnering with Infrastructure, Product Engineering, and Support teams to improve uptime, strengthen observability, establish SLOs, reduce operational toil through automation, and lead incident response initiatives. The ideal candidate brings experience supporting large-scale production environments and enjoys solving complex reliability challenges while influencing engineering practices across a rapidly growing organization. This is an opportunity to gain exposure to cutting-edge AI and GPU infrastructure, take ownership of high-impact initiatives, and help shape the reliability strategy of a platform relied upon by more than one million developers.
Required Skills & Experience:
5+ years of experience within major public cloud environment like AWS, GCPStrong Linux systems administration experience Strong networking fundamentals and troubleshooting skills Experience supporting containerized environments (Kubernetes preferred)Experience with monitoring, alerting, and observability tools Experience defining and managing SLIs, SLOs, and reliability metrics Incident response and postmortem experience Scripting or programming experience. Python, Go, Bash, or similar technologies Distributed systems and failure scenarios
Desired Skills & Experience:
Kubernetes Prometheus, Grafana, or similar monitoring platforms Experience supporting GPU infrastructure or AI/ML platforms Infrastructure as Code experience (Terraform preferred)CI/CD pipeline experience
What You Will Be Doing:
Tech Breakdown 40% Linux & Kubernetes Administration 25% Monitoring, Observability & Incident Response 20% Automation & Reliability Engineering 15% Distributed Systems &
Cloud Infrastructure Daily Responsibilities: 80% Hands-On Engineering 5%
Management Duties:
15% Team Collaboration The Offermedical, dental, and vision benefits Equity / Stock Options Remote equipment stipend•Annual learning and development budget Flexible PTOCareer Growth Within a Rapidly Scaling AI Infrastructure Company Applicants must be currently authorized to work in the US on a full-time basis now and in the future. Sponsorship is not available for this position#LI-JG2
Matching similar jobs
JOB OVERVIEW
Experience level
Senior
Location
Mt Laurel, NJ
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
8 days ago
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE