Salary: £100,000 – 100,000 per year
Requirements
- Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with the ability to implement these practices within an application or platform
- Fluency in at least one programming language, such as Python, Java Spring Boot, or .NET
- Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines
- Proficiency and hands-on experience in observability practices, including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk
- Proficiency in continuous integration and continuous delivery tools such as Jenkins, GitLab, or Terraform
- Experience with container technologies and container orchestration platforms such as ECS, Kubernetes, or Docker
- Experience troubleshooting common networking technologies and issues
- Ability to identify and resolve problems related to complex data structures and algorithms
- Ability to collaborate and communicate effectively across different levels and stakeholder groups
- Experience mentoring or coaching engineers on site reliability practices and engineering standards
- Familiarity with cloud platforms and infrastructure-as-code practices in large-scale enterprise environments
- Experience contributing to or leading communities of practice, internal knowledge sharing, or engineering guilds
- Exposure to chaos engineering or fault injection methodologies to proactively test system resilience
- Ability to evaluate and introduce emerging technologies that improve platform reliability and reduce operational toil
Responsibilities
- Demonstrate and champion site reliability culture and practices, and exert technical influence throughout our team
- Lead initiatives to improve the reliability and stability of our teams applications and platforms using data-driven analytics to improve service levels
- Collaborate with team members to identify comprehensive service level indicators and establish reasonable service level objectives and error budgets with customers
- Demonstrate a high level of technical expertise within one or more technical domains and proactively identify and solve technology-related bottlenecks in our areas of expertise
- Act as the main point of contact during major incidents for our application and identify and solve issues quickly to avoid financial losses
- Document and share knowledge within our organization via internal forums and communities of practice
- Take the lead on resiliency design reviews
- Break up complex problems into digestible work for other engineers
- Act as a technical lead for medium to large-sized products
- Provide advice and mentoring to other engineers
Technologies
- Cloud
- Datadog
- Docker
- Dynatrace
- GitLab
- Grafana
- Java
- Jenkins
- Kubernetes
- Marketing
- Prometheus
- Python
- Security
- Splunk
- Spring
- Spring Boot
- Terraform
- ASP.NET
- DevOps
last updated 36 week of 2026
#J-18808-Ljbffr…
