Site Reliability Engineer, Studios

Company: iMG world
Apply for the Site Reliability Engineer, Studios
Location: London
Job Description:

Overview

As Site Reliability Engineer at IMG, you design, build, and operate resilient platforms underpinning our digital, cloud, and live-broadcast services. You will improve reliability and observability while driving automation and disaster recovery readiness across on-prem and cloud environments. You collaborate with engineering and operations to raise release quality and incident response standards. You play a key role in scaling systems for live, business-critical workloads and supporting high-availability workflows.

Responsibilities

  • Design, build, and maintain reliable, scalable infrastructure across on-prem and cloud environments
  • Improve availability, latency, and efficiency through reliability engineering practices
  • Enhance observability with monitoring, logging, alerting, dashboards, and service health indicators
  • Define SLIs/SLOs, alerting standards, and runbooks for critical services
  • Automate provisioning, configuration, deployment, and recovery using IaC and scripting
  • Collaborate with software, platform, and broadcast engineering teams to improve resilience
  • Serve as escalation point for production incidents and drive post-incident follow-up
  • Lead root cause analysis and preventive actions for incidents
  • Support high availability, backup, failover, and disaster recovery design and testing
  • Enforce security, access control, patching, and best practices across infrastructure
  • Optimize capacity, cost, and performance; produce technical documentation
  • Support live events and critical operational workflows requiring rapid response and clear communication
  • Contribute to planning for new services, migrations, and platform enhancements with resilience in mind
  • Improve platform reliability, stability, and recovery across IMG services
  • Drive reduced mean time to detect/resolve incidents via observability and automation
  • Promote operational ownership and service standards across environments
  • Strengthen resilience for live client-facing workflows through tested failover approaches

Key requirements

  • Proven experience as Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar
  • Strong knowledge of Linux
  • Hands-on experience with AWS, Azure, or Google Cloud
  • Experience with Docker and Kubernetes
  • Experience with CI/CD tooling and modern software delivery
  • Hands-on with Infrastructure as Code tools like Terraform or CloudFormation
  • Experience with monitoring, logging, alerting, and observability design
  • Solid understanding of networking, security, architecture, and distributed systems
  • Scripting or programming in Python or Bash
  • Experience in high-availability, live production, or business-critical environments
  • Strong troubleshooting, calm decision-making under pressure, and continuous improvement mindset
  • Excellent communication and collaboration with technical and non-technical stakeholders
  • calm under pressure
  • collaboration
  • proactive ownership
  • Linux fundamentals
  • AWS/Azure/Google Cloud
  • Docker

Posted: September 14th, 2026