Overview
In this role you will own the stability and performance of the cloud platform within a transforming SaaS environment. You’ll work closely with cross-functional teams to restore services quickly after incidents and implement new tooling and processes to support the wider platform transformation. There are two tracks (Azure-focused and AWS-focused); you will drive reliability, automation, and observability while evolving toward leadership as the team grows.
Pay / Benefits
- bonus
- on call allowance
- competitive salary up to 75,000
- on-site presence (2 days per week)
- two role tracks (Azure and AWS)
- growth toward team lead position
Responsibilities
- Manage provisioning, configuration, and lifecycle automation of servers and networks using Terraform and Kubernetes tooling
- Lead incident resolution and drive improvements to MTTR and on-call efficiency
- Utilize observability platforms (Datadog, Prometheus, Grafana, Azure Monitor) to monitor application performance
- Implement and champion new ways of working and tooling across development and support teams
- Identify automation opportunities and implement changes to raise reliability across on-prem and Azure environments
- Participate in on-call rota and respond to incidents as needed
- Shape incident response processes to improve resilience across the platform
Key requirements
- SRE/DevOps or platform engineering experience in a customer-facing SaaS or MSP context
- Hands-on experience with Azure and on-premise VMs
- Infrastructure as Code experience with Terraform
- Container orchestration experience (Kubernetes or AKS)
- Experience with monitoring/observability tools (Prometheus, Grafana, Datadog, or Azure Monitor)
- Ability to implement new processes and tooling adopted by multi-team groups
- Strong networking skills across Azure and on-prem environments
- Windows or Linux systems administration (Linux preferred)
- Experience using AI tools to enhance efficiency
- Strong communication across internal teams and external customers
- ability to drive adoption of new ways of working
- problem-solving mindset
- Terraform
- Kubernetes/AKS
- Azure (Azure Monitor)
…
