Site Reliability Engineer

Company: The Business Connection Group
Apply for the Site Reliability Engineer
Location: London
Job Description:

Site Reliability Engineer | £130,000 – £150,000 | Remote

My client, a pioneering advanced technology organisation — a global leader in software‑defined networking platforms for the aerospace sector.

Their work underpins critical connectivity infrastructure and next‑generation space missions.

This is not a routine “maintain uptime” role — you will design the visibility layer that powers reliability for satellite constellations, ground station networks and deep‑space communications systems.

Key Responsibilities

  • Design, build and scale a unified observability platform (metrics, logging, tracing: Prometheus, Grafana, Loki, OpenTelemetry, Tempo/Jaeger)
  • Define and manage end-to-end SLOs, SLIs and error budgets to ensure production readiness and reliability
  • Partner with engineering to embed standards, set instrumentation best practices, and roll out consistent tooling
  • Automate deployment, scaling and lifecycle management via Infrastructure as Code (Terraform) and GitOps (ArgoCD)
  • Collaborate with infrastructure teams to deliver visibility across Kubernetes and multi-cloud environments
  • Lead monitoring, alerting and incident response; foster proactive reliability, blameless post‑mortems and continuous improvement

What You’ll Bring

  • 4+ years in SRE/reliability/platform engineering — focused on observability for large‑scale distributed systems
  • Hands‑on expertise building, scaling and running production observability stacks; diagnosing complex performance & availability issues
  • Strong production experience with GCP and Kubernetes
  • Practical IaC & GitOps experience for configuration and deployment management
  • Proficient in Go or Python for automation and tooling
  • Proven track record defining, implementing and governing SLO/SLI/error budget frameworks for high‑availability services

#J-18808-Ljbffr…

Posted: September 29th, 2026