Senior Site Reliability Engineer (LON)

Company: McNally Recruitment Ltd
Apply for the Senior Site Reliability Engineer (LON)
Location: London
Job Description:

Senior Site Reliability Engineer (London)

We’re working in collaboration to source a Senior Site Reliability Engineer for a large UK client. The role is mostly working remotely, with only 1 day per week being required to work in the London office.

  • In this key role, you’ll improve, drive, and embed non-functional and operational characteristics such as availability, performance, efficiency, change management, monitoring, security, incident response, and capacity planning of our products and services
  • You’ll enjoy significant stakeholder interaction, working in collaboration with engineers to ensure a principled approach to deliver change in a safe and secure way
  • This is a chance to join an inclusive team with a collaborative ethos and a commitment to innovation and professional development
  • You’ll work from home some of the time, but you’ll also spend a significant amount of time working from an office or hub

What you’ll do

  • Work closely with our feature team and other colleagues to meet defined service level objectives and continually improve systems and environments.
  • Define error budgets that support finding the right balance between risk and reliability.
  • Provide structure and help to our release process, suggesting and making improvements where possible.
  • Help scale systems sustainably through mechanisms like automation, evolving them by pushing for changes that improve reliability and velocity.
  • Coach and provide guidance to colleagues and the wider team, leading where required.

In addition to this, you’ll:

  • Proactively contribute new ideas and innovations to meet short-term and longer-term goals
  • Continually balance and manage any potential risks
  • Be accountable for the day-to-day health of both production and non-production environments and respond to any incidents as required
  • Provide technical expertise and input to establish the risk tolerance of products and services
  • Communicate incident status updates clearly and frequently to other teams, customers and stakeholders

The skills you’ll need

  • At least 10 years of hands-on experience, including as a Senior SRE with a proactive approach to spotting problems, areas for improvement, and performance bottlenecks.
  • Experience working with cloud-native microservices, including containerisation, management of Kubernetes workloads and API management.
  • Hands-on experience with Azure, Infrastructure as Code (IaC), and technologies such as PowerShell, JSON, Azure Bicep, ARM and Azure DevOps.
  • The client is moving to Terraform, which is essential, moving from Bicep (desirable).
  • Experience with Full Stack Observability using tools such as Grafana Stack, Log Analytics, AppInsights
  • Excellent knowledge of DevOps processes and principles
  • Knowledge of IT Service Management and automation of IT fulfilment processes through Orchestration and ServiceNow
  • Strong communication skills with the ability to proactively engage with a wide range of stakeholders

SALARY INCLUDES 10% Benefits-As-Cash

#J-18808-Ljbffr…

Posted: October 4th, 2026