Overview
In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. You will define the vision and roadmap for site reliability, drive best practices, and ensure resilience across production and non-production systems. You will manage people, budgets and schedules while maintaining strong technical standards and regulatory compliance. This is an opportunity to shape reliability at scale within a hybrid/London-based team, partnering across security, engineering and operations.
Responsibilities
- Define and own the vision and roadmap for site reliability and environment strategy
- Lead, mentor and develop a team of DevOps and environment engineers
- Set and enforce standards for environment provisioning, lifecycle management and patch governance
- Drive adoption of Infrastructure as Code and automation-first practices across environments
- Oversee monitoring, alerting and operational readiness to meet availability and performance objectives
- Partner with security, engineering and operations teams to reduce risk and ensure compliance
- Build continuous improvement processes and establish key performance metrics for reliability
Key requirements
- 8+ years of experience in site reliability, DevOps or platform engineering roles
- Strong knowledge of AWS platforms and cloud infrastructure principles
- Proven ability to manage teams and operational priorities in complex environments
- Experience implementing automation using Terraform or CloudFormation
- Experience with CI/CD pipelines and operational monitoring with tools such as CloudWatch, Prometheus and Grafana
- Understanding of OS patching for Windows Server and RHEL as part of governance practices
- Background in risk management, compliance frameworks and incident response processes
- Excellent leadership, communication and stakeholder management skills
- leadership
- stakeholder management
- communication
- Terraform
- CloudFormation
- CI/CD pipelines
…
