Staff Software Engineer – Databases SRE | UK | Remote

Company: Grafana Labs
Apply for the Staff Software Engineer – Databases SRE | UK | Remote
Location:
Job Description:

Overview

As Staff Software Engineer – SRE, you will own the reliability of Grafana Cloud databases (Mimir, Loki, Tempo, Pyroscope) delivered as a SaaS across AWS, GCP, and Azure. You’ll partner with embedded product engineering squads to meet high-SLA requirements and drive automation to scale reliability. You’ll define and evolve tenant-specific SLOs, lead incident response, and influence design for production scalability. This role offers a chance to shape resilient, scalable observability platforms for a global customer base.

Pay / Benefits

  • equity
  • bonus (if applicable)
  • remote work
  • 30 days annual leave
  • Grafana Shutdown Days for disconnect

Responsibilities

  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale reliability practices
  • Ensure customers meet SLO targets and evolve per-tenant SLOs
  • Proactively reduce SLO burn to prevent repeat incidents
  • Serve as primary escalation point and on-call for incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design docs and code reviews; influence feature design for scalability
  • Build automation to eliminate toil and improve observability and alerting
  • Collaborate with engineering leaders to define roadmaps and technical designs
  • Mentor engineers, advocate for SRE best practices in development

Key requirements

  • 8+ years of engineering experience, 4+ in SRE/production engineering
  • Formal customer reliability engineering experience preferred
  • Strong Kubernetes experience in AWS, GCP, or Azure
  • Familiarity with infrastructure-as-code tools (Helm, Terraform, Jsonnet)
  • Experience leading teams and mentoring engineers
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Proficiency in programming languages (Go, Python, Java, etc)
  • Knowledge of Linux internals, networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting abilities
  • Experience with blame-free incident response, PIRs, and post-mortems
  • Ability to reason about performance, scaling, and failure modes
  • Autonomy and self-direction within a collaborative engineering team
  • Comfortable partnering with product engineering teams
  • Curiosity, transparency, action bias, and kindness
  • problem-solving mindset
  • strong collaboration and teamwork
  • transparency and openness
  • Kubernetes (AWS/GCP/Azure)
  • Cloud platforms (AWS, GCP, Azure)
  • Infrastructure as code (Helm, Terraform, Jsonnet)

…

Posted: September 14th, 2026