Research Engineer, Pretraining Scaling – London

Company: Humanloop
Apply for the Research Engineer, Pretraining Scaling – London
Location: London
Job Description:

Overview

Join Anthropic’s ML Performance and Scaling team to keep production pretrained models training reliably at scale. You’ll operate at the research-engineering boundary, improving efficiency, observability, and reliability across the full training stack. Expect on-call incident response during launches and collaboration with global teams to push safe, steerable AI forward. This role offers hands-on work on large-scale ML systems with meaningful impact on our mission.

Pay / Benefits

  • competitive compensation
  • equity donation matching (optional)
  • generous vacation and parental leave
  • flexible working hours
  • office space for collaboration
  • visa sponsorship available

Responsibilities

  • Own production pretraining pipeline components (model operations, performance optimization, observability, reliability)
  • Debug complex issues across the stack—from hardware to training dynamics and evaluation infrastructure
  • Design and run experiments to improve training efficiency and model performance
  • Respond to on-call incidents during launches and coordinate cross-team solutions
  • Build/maintain production logging, monitoring dashboards, and evaluation infrastructure
  • Add capabilities to training codebase (e.g., long context support, new architectures)
  • Collaborate with SF and London teams plus Tokens, Architectures, and Systems groups
  • Document systems, debugging approaches, and lessons learned

Key requirements

  • Hands-on experience training large language models or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems
  • Enjoys both research and engineering work (ideally ~50/50 split)
  • Willing to be on-call for production systems during launches
  • Thrive in high-impact, dynamic environments and solving hard problems under pressure
  • Strong communication and cross-time-zone collaboration skills
  • Passion for AI safety and responsible scaling
  • communication
  • collaboration
  • problem-solving under pressure
  • JAX
  • TPU
  • PyTorch

…

Posted: October 1st, 2026