Overview
In this role you will co-develop reinforcement learning capabilities for large language models, balancing research exploration with robust engineering. You’ll help build scalable RL infrastructure, create test environments, and push model reasoning and tool-use capabilities. You work closely with researchers and engineers to advance safe, impactful AI systems at scale. This position offers opportunities to shape next-gen models and contribute to Anthropic’s mission of trustworthy AI.
Pay / Benefits
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Lovely office space for collaboration
Responsibilities
- Architect and optimize core RL infrastructure, including training abstractions and distributed experiment management across GPU clusters
- Scale systems to support increasingly complex research workflows
- Design and implement novel training environments and evaluation methodologies for RL agents
- Improve stack performance via profiling, optimization, and benchmarking
- Develop efficient caching and debugging for distributed training and evaluation
- Collaborate across research and engineering to build automated testing frameworks and scalable APIs
Key requirements
- Proficient in Python and async/concurrent programming (e.g., Trio)
- Experience with ML frameworks (PyTorch, TensorFlow, JAX)
- Industry experience in machine learning research
- Ability to balance research exploration with engineering implementation
- Enjoys pair programming and collaborative work
- Strong code quality, testing, and performance mindset
- Solid systems design and communication skills
- Passion for AI safety and beneficial AI
- Collaboration
- Communication
- Problem-solving
- Python
- async/concurrent programming (Trio)
- PyTorch
…
