Senior Site Reliability Engineer

Company: Elsevier
Apply for the Senior Site Reliability Engineer
Location: London
Job Description:

Overview

As a Senior Site Reliability Engineer at Elsevier, you will drive reliability, scalability, and performance of critical platforms. You’ll lead complex reliability initiatives, advance automation to reduce toil, and partner with engineering teams to fit real workflows. Your work strengthens service availability and operational readiness across AI-enabled services. You’ll shape incident response, observability, and platform automation to support mission-critical outcomes.

Pay / Benefits

  • Comprehensive Pension Plan
  • Generous vacation entitlement
  • Sabbatical leave option
  • Maternity, paternity, adoption and family care leave
  • Employee discounts
  • Internal communities and networks

Responsibilities

  • Create monitoring queries and establish service level baselines
  • Support senior engineers during incidents
  • Contribute to post-mortems and RCAs
  • Participate in disaster recovery tests
  • Implement automation and run code in production
  • Contribute to SRE knowledge documentation
  • Support deployment, monitoring, and reliability of AI-integrated services
  • Design infrastructure topology drawings and deployment workflows (with seniors)
  • Test availability, reliability, and recoverability in non-production environments
  • Benchmark and document test performance for production readiness reviews
  • Provide on-call support and incident automation scripts (failovers/rollbacks)
  • Create templated observability dashboards and code-based configurations
  • Influence SLOs and error budgets
  • Support migration to standard platforms and improve SDLC/CI/CD processes
  • Generate actionable reports on platform/product health

Key requirements

  • Terraform expertise (modules, providers, state management, drift, remote state)
  • Advanced AWS operations across ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, CloudWatch
  • GitHub Actions CI/CD experience (workflows, OIDC, approvals, runners, deployments)
  • ECS Fargate & containers knowledge (Docker, ECR, task definitions, IAM roles, health checks, autoscaling, ALB)
  • AWS networking and security (VPC, ALBs, Route53, TLS/ACM, IAM, OIDC, Secrets Manager, KMS)
  • Incident response and observability (logs, metrics, alarms, RCAs, runbooks)
  • Linux and automation (Bash/Python, AWS CLI automation, CI/CD tooling)
  • AI tooling deployment experience (production monitoring, reliability, security for AI features)
  • Developer enablement (support multiple teams, troubleshooting, documentation, self-service)
  • collaboration with cross-functional teams
  • problem-solving under pressure
  • clear communication of complex incidents
  • Terraform
  • AWS (ECS, RDS, S3, Lambda, DynamoDB, IAM, Route53, SQS, KMS, CloudWatch)
  • GitHub Actions CI/CD

…

Posted: October 1st, 2026