Senior Site Reliability Engineer – OpenTelemetry

Company: EPAM Systems
Apply for the Senior Site Reliability Engineer – OpenTelemetry
Location: London
Job Description:

Overview

In this role you will drive the design and expansion of a modern observability platform to boost monitoring, resilience, and operational stability for critical trading systems. You will work within the Production Engineering – Observability team to define standards, deliver enterprise-scale observability solutions, and promote best practices for operational excellence. The role involves hands-on instrumentation, data analytics, and dashboards to enable faster incident detection and reduced outages. This is a chance to shape observability at scale in a hybrid, finance-focused environment with a strong emphasis on automation and cross-team collaboration.

Responsibilities

  • Gather requirements and analyze current monitoring and observability setups
  • Define observability standards, telemetry strategies, and alerting frameworks
  • Implement OpenTelemetry instrumentation and OpenSearch-based solutions across apps and infrastructure
  • Design dashboards, analytics, and reporting to improve transparency and efficiency
  • Develop automation tools and processes to reduce manual overhead
  • Integrate observability frameworks with enterprise monitoring platforms like Geneos
  • Provide documentation and handover for long-term sustainability
  • Apply SRE principles to enhance stability, scalability and reduce incidents

Key requirements

  • Senior SRE or Observability Engineer with enterprise-scale implementation experience
  • Strong OpenTelemetry expertise including instrumentation and telemetry pipelines
  • Deep knowledge of OpenSearch for architecture, indexing, optimization, and analytics
  • Experience building dashboards and alerts with Grafana
  • Familiarity with enterprise monitoring tools such as Geneos
  • Practical understanding of SRE practices and automation
  • Background in high-availability or mission-critical environments; financial services experience is desirable
  • Knowledge of anomaly detection, alert correlation, and incident response automation
  • Exposure to observability in cloud-native or hybrid setups
  • collaboration
  • problem-solving
  • attention to detail
  • OpenTelemetry instrumentation
  • OpenTelemetry telemetry pipelines
  • OpenSearch data indexing and analytics

…

Posted: October 1st, 2026