White Circle in London is seeking a research engineer to build and maintain an internal benchmark suite spanning single/multi-turn content and agentic guardrails. You will collaborate with the core safety team to study agent behaviours and push evaluation tooling into production.
Applicants should have hands-on Python production experience, strong benchmark design skills, and a track record of shipping reliable code for ML systems. Hybrid work with offices in London and Paris is offered.
#J-18808-Ljbffr…
