Overview
Senior SRE role focused on keeping production cloud platforms observable, secure, scalable and reliable. You will lead automation and incident investigations within a growing Cloud Platform Engineering function, collaborating with Cloud Operations, Support, DevOps and Engineering teams. The role emphasizes Azure expertise, observability, and cost-conscious platform improvements to support mission-critical SaaS solutions. You will help shape architecture, governance and security practices while exploring AI-assisted tooling for automation and troubleshooting.
Pay / Benefits
- Bonus
- Medical Care
Responsibilities
- Protect and improve production environments as part of the SRE team
- Manage and prioritize a backlog of reliability, scalability and operational improvements
- Lead investigations into outages, performance issues and cloud expenditure
- Perform root-cause analysis and implement corrective actions
- Automate repetitive operational activities
- Provide technical leadership to Cloud Operations, Support, DevOps and Engineering teams
- Establish and maintain SLOs, SLAs, SLIs and error budgets
- Design and implement monitoring, alerting and dashboards across cloud and microservices
- Deploy observability tools (Grafana, Prometheus, Azure Monitor, OpenTelemetry)
- Develop custom metrics and queries for distributed microservices
- Create reusable Bicep/Terraform modules for monitoring and cloud infra
- Support and improve AKS-based Kubernetes environments
- Review and optimize platform performance, security and cost
- Contribute to cloud architecture and scalable platform solutions
- Support continuous improvement of deployment and provisioning processes
- Ensure security, governance and compliance requirements are met
- Explore AI-assisted tooling to enhance automation and productivity
Key requirements
- Six+ years in SRE/related cloud role
- Strong hands-on Azure expertise
- Production experience with Kubernetes/AKS
- Extensive experience in observability and monitoring
- Experience with Grafana, Prometheus, OpenTelemetry, Elasticsearch
- Custom metrics, queries, dashboards for microservices
- Scripting or development skills (PowerShell, Python, C#)
- IaC experience (Bicep, ARM, Terraform)
- Git or other VCS
- Knowledge of SQL Server, Elasticsearch, YAML/JSON/XML
- Understanding of microservices, cloud platforms and containerisation
- Experience defining/woking with SLOs, SLAs, SLIs and error budgets
- Troubleshooting and root-cause analysis skills
- Security, governance and compliance awareness
- Experience in transformation projects and live-service environments
- Technical leadership
- Collaborative cross-functional communication
- Problem-solving and analytical mindset
- Microsoft Azure
- Kubernetes / AKS
- Grafana
…
