Overview
In this role you will help sustain and strengthen large-scale cloud infrastructure for essential national services. You’ll operate in a 24/7 NOC with a focus on proactive reliability, automation, and rapid incident resolution. You’ll work with cross-functional teams to improve monitoring, observability, and platform stability. The opportunity offers a challenging, hands-on environment where engineering-driven practices shape operations and cloud evolution.
Pay / Benefits
- remote working (UK-based)
- competitive salary
- bonus
- strong benefits package
Responsibilities
- Monitor business-critical AWS production environments for performance and uptime
- Own live incidents in a 24/7 rotating shift model and drive resolution
- Investigate and resolve complex technical issues to restore services quickly
- Develop automation to replace repetitive manual tasks and improve robustness
- Enhance monitoring, alerting, and observability tooling across the cloud estate
- Collaborate with software, platform, cloud and security teams to raise reliability and standards
- Participate in incident retrospectives and implement lessons learned
- Manage containerised applications on Terraform and Docker
Key requirements
- Production engineering, cloud operations, or NOC background with hands-on AWS experience
- Proficiency with Linux administration
- Experience with Terraform and Docker
- Incident handling in live production environments
- Scripting in Python, Bash or Go
- Experience with observability/monitoring tools (Grafana, Prometheus, Datadog, Splunk, CloudWatch)
- Strong networking knowledge (DNS, TCP/IP, load balancing)
- Drive to automate and improve operational excellence
- Problem-solving under pressure
- Cross-functional collaboration
- Proactive and detail-oriented
- AWS infrastructure
- Terraform
- Docker
…
