Lead Site Reliability Engineer – Glasgow

Company: Hackajob Ltd
Apply for the Lead Site Reliability Engineer – Glasgow
Location: Glasgow
Job Description:

Salary: £100,000 – 100,000 per year

Requirements

  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with the ability to implement these practices within an application or platform
  • Fluency in at least one programming language, such as Python, Java Spring Boot, or .NET
  • Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines
  • Proficiency and hands-on experience in observability practices, including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk
  • Proficiency in continuous integration and continuous delivery tools such as Jenkins, GitLab, or Terraform
  • Experience with container technologies and container orchestration platforms such as ECS, Kubernetes, or Docker
  • Experience troubleshooting common networking technologies and issues
  • Ability to identify and resolve problems related to complex data structures and algorithms
  • Ability to collaborate and communicate effectively across different levels and stakeholder groups
  • Experience mentoring or coaching engineers on site reliability practices and engineering standards
  • Familiarity with cloud platforms and infrastructure-as-code practices in large-scale enterprise environments
  • Experience contributing to or leading communities of practice, internal knowledge sharing, or engineering guilds
  • Exposure to chaos engineering or fault injection methodologies to proactively test system resilience
  • Ability to evaluate and introduce emerging technologies that improve platform reliability and reduce operational toil

Responsibilities

  • Demonstrate and champion site reliability culture and practices, and exert technical influence throughout our team
  • Lead initiatives to improve the reliability and stability of our teams applications and platforms using data-driven analytics to improve service levels
  • Collaborate with team members to identify comprehensive service level indicators and establish reasonable service level objectives and error budgets with customers
  • Demonstrate a high level of technical expertise within one or more technical domains and proactively identify and solve technology-related bottlenecks in our areas of expertise
  • Act as the main point of contact during major incidents for our application and identify and solve issues quickly to avoid financial losses
  • Document and share knowledge within our organization via internal forums and communities of practice
  • Take the lead on resiliency design reviews
  • Break up complex problems into digestible work for other engineers
  • Act as a technical lead for medium to large-sized products
  • Provide advice and mentoring to other engineers

Technologies

  • Cloud
  • Datadog
  • Docker
  • Dynatrace
  • GitLab
  • Grafana
  • Java
  • Jenkins
  • Kubernetes
  • Marketing
  • Prometheus
  • Python
  • Security
  • Splunk
  • Spring
  • Spring Boot
  • Terraform
  • ASP.NET
  • DevOps

last updated 36 week of 2026

#J-18808-Ljbffr…

Posted: September 9th, 2026