Senior Site Reliability Engineer
radley jamesNew York, NY
Senior Site Reliability Engineer
radley jamesNew York, NY
yesterday
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems AdministratorsIndustries
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesOther Computer Related ServicesAbout the role
A leading High-Frequency Trading firm is seeking an experienced Senior Site Reliability Engineer. This is a critical role where reliability, performance, automation and operational excellence are essential. You will work closely with software engineers, infrastructure teams and traders to build and maintain highly available, resilient and scalable platforms supporting time-sensitive trading systems. The ideal candidate will have a strong background in Site Reliability Engineering, DevOps and cloud infrastructure, with significant hands-on experience working with AWS.
The Role:
As a Senior SRE, you will be responsible for ensuring the reliability, scalability and performance of the technology platforms.
You will:
Design, build and maintain highly available and resilient infrastructure on AWS.Develop and improve automation across infrastructure, deployment and operational processes.
Establish and maintain monitoring, observability, alerting and incident-management practices.
Work closely with development teams to improve application reliability and performance.
Participate in the design and implementation of highly resilient systems supporting trading and business-critical workloads.
Lead incident response, troubleshooting and root-cause analysis for complex production issues.
Identify and eliminate recurring operational problems through automation and engineering.
Contribute to capacity planning, performance optimisation and disaster-recovery strategies.
Improve CI/CD pipelines and deployment processes.
Champion SRE and DevOps best practices across the engineering organisation.
Mentor engineers and provide technical leadership on reliability and infrastructure matters.
Essential Experience:
We are looking for candidates with:5+ years' experience in SRE, DevOps, Platform Engineering or a similar infrastructure-focused role.
Strong, hands-on AWS experience in production.
Excellent understanding of AWS services such as EC2, EKS/ECS, S3, IAM, VPC, CloudWatch, RDS and Route 53.Strong Linux/Unix administration and troubleshooting skills.
Experience with Infrastructure as Code, ideally Terraform.
Strong scripting/programming ability with Python.
Experience with Kubernetes and containerised environments.
Strong understanding of CI/CD principles and tooling.
A strong understanding of networking, security and distributed systems.
Proven experience managing production incidents and conducting root-cause analysis.
Experience building systems with high availability, resilience and fault tolerance in mind.
Matching similar jobs
JOB OVERVIEW
Experience level
Senior
Location
New York, NY
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
yesterday
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE