Overview
As SRE Engineering Manager, you will lead reliability-focused initiatives across engineering and LiveOps to deliver scalable, secure, cloud-native SaaS for education, health and care, and local government. You’ll partner with Platform, Security, and Product teams to raise availability and performance while embedding observability and incident response into delivery. You’ll drive automation, IaC, and SRE practices, mentoring engineers and shaping a culture of reliability. This is a high-impact role at a GovTech leader shaping services used by millions.
Pay / Benefits
- 25 Days Annual Leave + bank holidays
- Pension Contributions – 5% employer match
- Income protection – up to 75% salary
- Life Assurance – 4x salary
- Private Medical Insurance
- Health Cash Plan
Responsibilities
- Deliver reliability improvements with SRE practices across teams, improving availability, performance, and incident response
- Bridge engineering and operations by collaborating with Platform and LiveOps to scale securely and efficiently
- Instrument services with observability (metrics, tracing, logging) to provide shared visibility
- Lead incident response, run blameless post-mortems, and feed lessons into engineering and product planning
- Drive automation and standardisation through IaC and CI/CD (Terraform, GitHub Actions, Ansible, Packer, Kubernetes)
- Establish and evolve core SRE practices (SLOs, on-call, toil reduction, production readiness) and tailor them to organisational maturity
- Provide hands-on technical leadership: review architectures, prototype reliability solutions, troubleshoot complex issues, and make trade-off decisions
- Identify systemic reliability risks and manage technical debt to safeguard service quality
- Mentor engineers, foster psychological safety, and drive accountability within the team
- Collaborate with Security and Compliance to embed resilience and regulatory requirements (SOC2, ISO27001, PCI) in design and delivery
Key requirements
- Experience leading SRE, DevOps, or Cloud/Software engineering teams (5+ engineers) in a SaaS or multi-product environment
- Strong cloud architecture background (AWS, Azure, GCP) and hybrid/cloud migration experience
- Hands-on ability to read and review application/infrastructure code and make architecture decisions
- Proven background in observability tooling and incident management, using data to diagnose production problems
- Excellent communication and facilitation across Product, Platform, Security, Support and LiveOps
- Mentorship skills to foster clarity, alignment, and confidence in engineers
- Passion for reliability and user experience as a cultural value, embedding SRE principles across teams
- communication
- facilitation
- mentorship
- AWS/Azure/GCP cloud architecture
- IaC (Terraform, Ansible, Packer)
- Kubernetes
…
