Site Reliability Engineer

Company: Trainline
Apply for the Site Reliability Engineer
Location: London
Job Description:

Overview

Join Trainline as a mid-level Site Reliability Engineer within ReliabilityOps. You’ll help keep a cloud-first platform observable, scalable, and resilient, partnering with product teams to enable safe delivery. You’ll handle incident response, post-incident reviews, and tooling to improve MTTR and SLO adherence. This role offers hands-on architecture influence and shared ownership in a fast-growing fintech-like travel platform.

Pay / Benefits

  • private healthcare & dental insurance
  • work from abroad policy
  • 2-for-1 share purchase plans
  • EV Scheme to reduce carbon emissions
  • extra festive time off
  • family-friendly benefits

Responsibilities

  • Develop understanding of system architecture, dependencies, and failure modes across the platform
  • Participate in production incident response, investigations, mitigation, and service restoration
  • Contribute to post-incident reviews and follow-up actions for reliability, scalability, and resilience
  • Take part in on-call rotation
  • Design, build, and maintain observability using metrics, logs, events, and traces
  • Improve monitoring and alerting aligned with business impact to reduce noise and MTTD
  • Surface operational data quickly during live incidents
  • Make informed tooling and technology choices using SRE principles
  • Support AWS-hosted infrastructure and shared platform services using IaC and CI/CD tooling
  • Collaborate with product engineering to ensure services are deployment-ready and operationally safe
  • Advise on reliability and resilience practices
  • Write and maintain reliable, well-structured code and scripts
  • Prioritise work effectively using agile processes
  • Contribute to broader reliability discussions and ongoing improvements

Key requirements

  • SRE concepts such as SLI, SLO and error budgets
  • Hands-on observability tooling experience (New Relic, ELK, Influx, Grafana)
  • Experience with AWS or similar cloud providers
  • Troubleshooting Linux operating systems
  • Scripting in at least one language (preferably Python)
  • Understanding of load balancing, reverse proxy concepts, upstream config/health checks
  • Knowledge of application architecture concepts (threading, queues, readiness/health checks, circuit breakers, backoff, throttling)
  • Experience with time series data management (retention, cardinality, moving averages)
  • Experience with GitHub Actions and Terraform
  • Strong collaboration and agile workflow
  • collaboration
  • growth mindset
  • ability to challenge and be challenged
  • SRE fundamentals (SLI/SLO, error budgets)
  • Observability tooling: New Relic, Elastic/ELK, Grafana, Influx
  • AWS and cloud-native infrastructure

…

Posted: September 30th, 2026