Overview
In this role you will advance AGI safety and alignment research, translating methods into production-ready systems. You will collaborate with cross-functional teams to ensure research is adopted in products, and explore interpretability to understand AI decisions. You will shape defense-in-depth controls for AGI deployments, advise leadership on risk, and contribute to scalable alignment efforts for frontier models.
Responsibilities
- Research new alignment methods for frontier models and study alignment failures
- Develop AGI control systems and implement them in production
- Investigate interpretability techniques to understand AI system thinking
- Collaborate with product teams to ensure research is adopted in products
- Advise executive leadership on safety risks using frontier safety framework
- Contribute to alignment stress testing, evaluative research, and tool development
- Explore training techniques like debate for aligning superhuman AI
- Support safety research with tools for model forensics and eval awareness
- Participate in cross-functional efforts to reduce existential and catastrophic AI risk
Key requirements
- Bachelor’s degree in Computer Science, related Software Engineering field, or equivalent practical experience
- 3 years of experience in software development, ML engineering, or ML research
- Experience working with research teams
- collaboration with cross-functional teams
- clear technical communication
- problem solving under high risk contexts
- ML research and engineering experience
- training large models (supervised finetuning, RLHF)
- interpretability research
…
