AI Security Institute in London is seeking a Research Engineer to join the Human Influence team. The role blends research and engineering, focusing on scalable post-training and fine-tuning of LLMs, with reinforcement learning and validation on large compute clusters.
You will design RL environments to curb deceptive interactions, apply interpretability methods to reveal model risks, and build the systems that deliver repeatable evaluations and benchmarks.
#J-18808-Ljbffr…
