Overview
Join Anthropic’s ML Performance and Scaling team to keep production pretrained models training reliably at scale. You’ll operate at the research-engineering boundary, improving efficiency, observability, and reliability across the full training stack. Expect on-call incident response during launches and collaboration with global teams to push safe, steerable AI forward. This role offers hands-on work on large-scale ML systems with meaningful impact on our mission.
Pay / Benefits
- competitive compensation
- equity donation matching (optional)
- generous vacation and parental leave
- flexible working hours
- office space for collaboration
- visa sponsorship available
Responsibilities
- Own production pretraining pipeline components (model operations, performance optimization, observability, reliability)
- Debug complex issues across the stack—from hardware to training dynamics and evaluation infrastructure
- Design and run experiments to improve training efficiency and model performance
- Respond to on-call incidents during launches and coordinate cross-team solutions
- Build/maintain production logging, monitoring dashboards, and evaluation infrastructure
- Add capabilities to training codebase (e.g., long context support, new architectures)
- Collaborate with SF and London teams plus Tokens, Architectures, and Systems groups
- Document systems, debugging approaches, and lessons learned
Key requirements
- Hands-on experience training large language models or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems
- Enjoys both research and engineering work (ideally ~50/50 split)
- Willing to be on-call for production systems during launches
- Thrive in high-impact, dynamic environments and solving hard problems under pressure
- Strong communication and cross-time-zone collaboration skills
- Passion for AI safety and responsible scaling
- communication
- collaboration
- problem-solving under pressure
- JAX
- TPU
- PyTorch
…
