Site Reliability Engineering (SRE) / Observability Technical Lead

Company: NTT DATA
Apply for the Site Reliability Engineering (SRE) / Observability Technical Lead
Location: London
Job Description:

Overview

As a Site Reliability Engineer (SRE) and Observability Technical Lead, you will shape and execute our observability and reliability strategy across clients. You’ll design and implement monitoring, tracing, and logging frameworks to improve system performance, reliability, and scalability. You will lead cross-functional teams, mentor juniors, and partner with product, sales, and vendors to drive solution design. This role offers impact across cloud, CI/CD, and IaC initiatives, with a strong focus on automation and best practices.

Pay / Benefits

  • flexible work options
  • learning and development opportunities
  • focus on wellbeing
  • inclusive culture
  • equal opportunities employer
  • disability confident commitment

Responsibilities

  • Lead development and management of observability and reliability frameworks across the organization
  • Design and implement monitoring and observability standards with engineering teams
  • Manage IaC initiatives using Terraform and coordinate with cloud/infrastructure teams
  • Drive automation for monitoring, alerting, and logging pipelines
  • Develop and maintain observability roadmaps for tracing, logging, and metrics
  • Collaborate with product management, sales, and pre-sales for technical guidance
  • Enhance CI/CD pipelines and deployment reliability with observability integration
  • Engage with vendors/partners to select and integrate monitoring solutions
  • Mentor and develop junior engineers and analysts in reliability and observability

Key requirements

  • 5+ years in SRE, Observability, or DevOps with leadership experience
  • Proven expertise with APM tools (New Relic, Datadog, AppDynamics, Dynatrace)
  • Hands-on OpenTelemetry for distributed tracing
  • Strong IaC experience with Terraform
  • Cloud platform experience (AWS, GCP, or Azure)
  • Automation/configuration management experience (Ansible, Chef, Puppet)
  • Deep knowledge of CI/CD (GitHub Actions, Jenkins, Azure DevOps)
  • Kubernetes/containerized environments (Docker, Helm)
  • Familiarity with ELK Stack or Splunk
  • Excellent leadership and communication skills
  • leadership
  • communication
  • collaboration
  • OpenTelemetry
  • Terraform
  • AWS/GCP/Azure cloud platforms

…

Posted: September 30th, 2026