Lead Site Reliability Engineer

Company: Spectrum IT Recruitment
Apply for the Lead Site Reliability Engineer
Location: Southampton
Job Description:

Overview

Senior SRE role focused on keeping production cloud platforms observable, secure, scalable and reliable. You will lead automation and incident investigations within a growing Cloud Platform Engineering function, collaborating with Cloud Operations, Support, DevOps and Engineering teams. The role emphasizes Azure expertise, observability, and cost-conscious platform improvements to support mission-critical SaaS solutions. You will help shape architecture, governance and security practices while exploring AI-assisted tooling for automation and troubleshooting.

Pay / Benefits

  • Bonus
  • Medical Care

Responsibilities

  • Protect and improve production environments as part of the SRE team
  • Manage and prioritize a backlog of reliability, scalability and operational improvements
  • Lead investigations into outages, performance issues and cloud expenditure
  • Perform root-cause analysis and implement corrective actions
  • Automate repetitive operational activities
  • Provide technical leadership to Cloud Operations, Support, DevOps and Engineering teams
  • Establish and maintain SLOs, SLAs, SLIs and error budgets
  • Design and implement monitoring, alerting and dashboards across cloud and microservices
  • Deploy observability tools (Grafana, Prometheus, Azure Monitor, OpenTelemetry)
  • Develop custom metrics and queries for distributed microservices
  • Create reusable Bicep/Terraform modules for monitoring and cloud infra
  • Support and improve AKS-based Kubernetes environments
  • Review and optimize platform performance, security and cost
  • Contribute to cloud architecture and scalable platform solutions
  • Support continuous improvement of deployment and provisioning processes
  • Ensure security, governance and compliance requirements are met
  • Explore AI-assisted tooling to enhance automation and productivity

Key requirements

  • Six+ years in SRE/related cloud role
  • Strong hands-on Azure expertise
  • Production experience with Kubernetes/AKS
  • Extensive experience in observability and monitoring
  • Experience with Grafana, Prometheus, OpenTelemetry, Elasticsearch
  • Custom metrics, queries, dashboards for microservices
  • Scripting or development skills (PowerShell, Python, C#)
  • IaC experience (Bicep, ARM, Terraform)
  • Git or other VCS
  • Knowledge of SQL Server, Elasticsearch, YAML/JSON/XML
  • Understanding of microservices, cloud platforms and containerisation
  • Experience defining/woking with SLOs, SLAs, SLIs and error budgets
  • Troubleshooting and root-cause analysis skills
  • Security, governance and compliance awareness
  • Experience in transformation projects and live-service environments
  • Technical leadership
  • Collaborative cross-functional communication
  • Problem-solving and analytical mindset
  • Microsoft Azure
  • Kubernetes / AKS
  • Grafana

…

Posted: October 3rd, 2026