Lead Site Reliability / DevOps Engineer

Company: JP Morgan Chase
Apply for the Lead Site Reliability / DevOps Engineer
Location: Glasgow
Job Description:

Overview

In this Lead SRE role, you guide reliability and performance across large-scale applications in a global bank. You’ll shape resiliency practices, mentor engineers, and lead incident management to prevent financial impact. You will drive data-driven improvements, align service levels with stakeholders, and advance OpenTelemetry-based observability across hybrid environments. This position offers impact across platforms and teams, with a focus on scalable, secure, and stable services that enable business excellence.

Responsibilities

  • Promote site reliability culture and exert technical influence across the team
  • Lead reliability initiatives and use analytics to improve service levels
  • Collaborate to define service level indicators and objectives with stakeholders
  • Provide expert guidance to solve bottlenecks in key technical domains
  • Serve as incident commander for major outages and drive rapid resolution
  • Document and share knowledge within internal communities
  • Design, implement, and maintain OpenTelemetry pipelines for large-scale observability
  • Support telemetry ingestion, processing, and export to backends (InfluxDB, Prometheus, Elasticsearch, OpenSearch) for performance, monitoring, logging and alerting
  • Refactor legacy telemetry code toward standardized OpenTelemetry instrumentation to reduce technical debt while preserving stability

Key requirements

  • Formal training or certification in software engineering concepts with advanced hands-on experience
  • Deep proficiency in reliability, scalability, performance, security, and enterprise architecture
  • Fluency in at least one programming language (Java, Python, Go, etc.)
  • Strong observability experience with tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc.
  • Experience with CI/CD tools (Jenkins, GitLab, Terraform, etc.)
  • Experience with containers and orchestration (ECS, Kubernetes, Docker, etc.)
  • Hands-on experience with OpenTelemetry collectors in production, including OTLP endpoints and receivers
  • Ability to collaborate across levels and stakeholder groups
  • collaboration
  • mentoring
  • strong communication
  • observability and monitoring
  • SLO/SLI/ error budgeting
  • OpenTelemetry instrumentation and collectors

Posted: September 14th, 2026