About the role

To support enterprise AI platforms, the contract Platform Reliability Engineer will build observability, implement OpenTelemetry instrumentation, and manage cloud cost monitoring while working remotely.
Key responsibilities: Design and implement OpenTelemetry-based instrumentation and centralized logging across AI agents and services Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability dashboards for platform components Configure proactive alerts for service degradation, execution failures, and resource utilization, supporting incident management and operational runbooksRequired qualifications 7+ years of experience in platform engineering, SRE, DevOps, or production infrastructure operations Hands-on experience with OpenTelemetry SDKs and telemetry pipelines Strong experience in implementing distributed tracing, metrics, logging, and alerting Experience monitoring AWS infrastructure and services, including CloudWatch and EKS Proficiency in Python or another scripting language, with experience in Terraform or equivalent IaC tools

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

New York, NY

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

9 days ago

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE