Staff Software Engineer, AI Reliability Engineering

Company: Humanloop
Apply for the Staff Software Engineer, AI Reliability Engineering
Location: London
Job Description:

Overview

In this role you will improve the reliability and scalability of Anthropic’s AI serving systems. You’ll collaborate across teams to design robust, observable infrastructure that supports millions of external customers and high-traffic workloads. You will lead incident response and drive continuous improvements to ensure a consistently excellent customer experience. This role offers the chance to shape reliability playbooks in a fast-growing, mission-driven AI company.

Pay / Benefits

  • competitive compensation and benefits
  • optional equity donation matching
  • generous vacation and parental leave
  • flexible working hours
  • office space for collaboration

Responsibilities

  • Define Service Level Objectives for large language model serving and training systems balancing availability and velocity
  • Design and implement monitoring for availability, latency, and other key metrics
  • Build high-availability model serving infrastructure for millions of users and internal workloads
  • Develop automated failover and recovery across multi-region, multi-cloud deployments
  • Lead incident response for critical AI services and drive systematic post-incident improvements
  • Create cost optimization systems focusing on accelerator utilization (GPU/TPU/Trainium) and efficiency

Key requirements

  • Extensive experience with distributed systems observability and monitoring at scale
  • Experience operating AI infrastructure including model serving, batch inference, and training pipelines
  • Proven track record implementing SLO/SLA frameworks for business-critical services
  • Ability to work with traditional metrics (latency, availability) and AI metrics (model performance, training convergence)
  • Experience with chaos engineering and resilience testing
  • Strong communication skills and ability to bridge ML engineers with infrastructure teams
  • strong communication
  • cross-functional collaboration
  • problem-solving mindset
  • distributed systems observability
  • SLO/SLA framework implementation
  • AI infrastructure for serving and training pipelines

…

Posted: September 14th, 2026