SRE | Permanent | London, Hybrid, AWS

Company: Source Group International
Apply for the SRE | Permanent | London, Hybrid, AWS
Location: London
Job Description:

Overview

As a Site Reliability Engineer, you will shape reliable, observable, secure and cost-efficient systems on AWS. You collaborate with development, platform, and incident teams to define measurable reliability metrics and implement tooling to improve speed, stability, and scalability. You’ll balance performance with cost, contribute to pre-deploy checks and incident response, and drive resilience through automation and chaos experiments. This role offers impact on high‑volume gaming platforms and data‑driven architecture.

Responsibilities

  • Define, measure, and manage SLOs/SLIs with error budgets to guide delivery decisions
  • Enhance observability across services using metrics, logs, and traces
  • Lead cost optimisation: monitor spend, right-size workloads, tune autoscaling
  • Improve production readiness via pre-deployment checks and post-release validation
  • Introduce and run chaos engineering experiments to strengthen resilience
  • Automate operational processes to reduce toil across the stack
  • Support major incident response, root-cause analysis, and continual improvement
  • Collaborate cross-functionally to raise standards for stability, security, performance, and compliance

Key requirements

  • 3+ years in SRE, Platform, or DevOps roles in production
  • Strong Kubernetes operational experience (on-prem and AWS EKS)
  • Hands-on experience defining and operating SLOs/SLIs, alerting, and incident workflows
  • Deep understanding of observability and telemetry (monitoring, logging, tracing)
  • Infrastructure as Code with Terraform; GitOps and CI/CD experience
  • Scripting in Python, Bash, or Go
  • Ability to balance cost efficiency with reliability and performance
  • Excellent communication across multiple teams
  • cross-functional collaboration
  • strong communication
  • problem-solving mindset
  • Kubernetes (on-prem and AWS EKS)
  • SLOs/SLIs and alerting
  • Observability: metrics, logs, traces

Posted: September 14th, 2026