Overview
In this role, you will lead cutting-edge reinforcement learning research and post-training at scale to improve DeepL’s multi-modal models. You’ll shape how models learn beyond pre-training, aiming for safer, more capable systems used in production. You’ll collaborate with cross-functional teams and external partners to advance the foundation model stack. This is a research-driven position at the Foundation Model Task Adaptation team, offering impact through real-world deployment and scalable experiments. Join us to help push the boundaries of AI alignment and practical AI deployment at global scale.
Pay / Benefits
- Hybrid work
- 30 days annual leave
- Hack Fridays monthly
- Diverse, international team
- Mental health resources
- Competitive benefits tailored to location
Responsibilities
- Design, implement, and deploy reinforcement learning research and post-training pipelines at scale
- Post-train large multi-modal models to align with human intent and enhance capabilities, safety, and efficiency
- Manage end-to-end research lifecycle from ideation to production deployment
- Foster external collaborations with academic and industrial partners
- Maintain rigorous experimentation, reproducibility, and model evaluation standards
- Collaborate with Engineering, ML Platform, and HPC teams to deliver robust model updates to users
Key requirements
- Deep technical background in reinforcement learning or large-scale model alignment to production
- Strong leadership and track record of self-directed research delivering tangible results
- Solid mathematical background; advanced degree or equivalent in relevant field
- Proficiency in Python and at least one ML framework (PyTorch, TensorFlow, or JAX)
- Experience with large compute clusters and ML infrastructure a plus
- Experience scaling and deploying LLMs or foundation models in real-world systems a plus
- Expertise in RLHF/RLAIF/RLVR is a plus
- Leadership
- Creative problem solving
- Strong collaboration and communication
- Python
- PyTorch
- TensorFlow
…
