Software Engineer, Infrastructure, AI Labs
epiq systemsBrooklyn, NY
yesterday
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems AdministratorsIndustries
Computer Systems Design ServicesSoftware PublishersCustom Computer Programming ServicesAbout the role
Overview
In this role you’ll build and run the cloud infrastructure that powers Epiq AI Labs’ AI platform. You’ll own infrastructure from design to production, focusing on reliability, scalability, security, and observability. You’ll collaborate with backend, AI, product, and security teams to scale deployment, monitoring, and compliance. You’ll work in a fast-moving, collaborative environment that treats infrastructure as a product and aims for end-to-end ownership. This is a chance to shape a cutting-edge platform at scale.
Compensation / Benefitsenterprise-wide learningmobility programsflexible work arrangementsgrowth opportunities
Responsibilities Design and implement cloud infrastructure with Terraform, including modular components and drift detection Operate production Kubernetes clusters with autoscaling, resource governance, network policy, and lifecycle management Build and maintain CI/CD and release pipelines with progressive delivery and rollback capabilities Define platform service-level objectives and establish metrics, tracing, alerting, and error budgets Implement security and compliance infrastructure, including hardening, audit logging, data residency controls, retention, and audit evidence collection Manage networks, secrets, keys, credentials, TLS, and certificate lifecycles Contribute to incident response, post-incident reviews, developer tooling, design docs, and platform readiness
Key requirements 3+ years in infrastructure/platform/SRE roles Experience operating production infrastructure with on-call responsibility Hands-on with major cloud platforms (AWS, GCP, or Azure)IaC experience with Terraform and modular design Production Kubernetes experience with scaling, upgrades, resource limits, and troubleshooting Ownership of CI/CD pipelines (Git Hub Actions, Azure Dev Ops, or similar)Observability tooling experience (Prometheus, Grafana, or Open Telemetry) with SLOsIncident-command experience in production Secrets and certificate management at scale Python, Go, or similar for tooling/automation Strong system-design and architectural documentationcollaborative mindsetstrong communicationownership and accountability Terraform (modules, drift detection)Kubernetes production operationsCI/CD tooling (Git Hub Actions, Azure Dev Ops)
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Brooklyn, NY
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
yesterday
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE