Lead Site Reliability Engineer

Company: JP Morgan Chase
Apply for the Lead Site Reliability Engineer
Location: London
Job Description:

Overview

In this Lead SRE role, you drive reliability and resilience across applications and platforms within the Infrastructure Platforms team. You lead incident management, mentor engineers, and steer data-driven efforts to meet service levels. You will shape AI-assisted reliability workflows and guide teams through complex problems, contributing to scalable, secure systems that support the firm’s objectives. This is a high-impact, leadership-focused opportunity in a globally recognized organization.

Responsibilities

  • Champions SRE culture and exerts technical influence across the team
  • Leads initiatives to improve reliability using data-driven analytics to enhance service levels
  • Collaborates to define service level indicators and establish SLOs and error budgets with stakeholders
  • Demonstrates deep technical expertise to resolve bottlenecks in assigned domains
  • Acts as point of contact during major incidents to accelerate resolution and reduce financial impact
  • Documents and shares knowledge internally via forums and communities of practice
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices with traceability, resiliency, and security controls

Key requirements

  • Formal training or certification on SRE concepts and advanced applied experience
  • Design and code complex problems in public cloud like AWS
  • Deep proficiency in reliability, scalability, performance, security, enterprise architecture, toil reduction
  • Fluency in Python and knowledge of software processes with depth in technical disciplines
  • Proficiency and experience in observability using Grafana, Dynatrace, Prometheus, Datadog, Splunk
  • Proficiency in CI/CD tools (Jenkins, GitLab, Terraform) and container orchestration (ECS, Kubernetes, Docker)
  • Experience troubleshooting networking technologies
  • Ability to solve problems related to complex data structures and algorithms
  • Commitment to self-education, teaching new languages, and collaborating across stakeholder groups
  • Experience using enterprise-authorized AI capabilities to improve SRE workflows with proper validation and guardrails
  • leadership
  • mentoring
  • collaboration
  • AWS and cloud design
  • Python development
  • observability and monitoring tools (Grafana, Dynatrace, Prometheus, Datadog, Splunk)

…

Posted: September 30th, 2026