- We’re looking for a senior engineer to lead the development of our flagship product, Inspect Evals .
- The successful candidate will join the founding team of a well-resourced, fast moving nonprofit start-up, at the centre of the AI safety ecosystem.
- This role is an excellent fit for user-focused SWEs with a passion for scientific software, data and research
About Generality Labs:
- Generality Labs is an AI safety R&D lab based in London.
- We were founded in 2026, spinning off from Arcadia Impact with $4.4M in seed funding from a US donor
- We’re a team of 2 cofounders, and 5+ contractors who have previously worked at places like METR, CERN and Anthropic.
- Our mission is to build tools which address key scientific challenges in AI evaluation, and then drive adoption of these solutions where they are most needed – governments, frontier labs, 3rd party AI evaluators.
- Our ultimate goal is to ensure the field can accurately evaluate the safety of advanced AI (which is currently not the case).
About Inspect Evals:
Inspect Evals was founded in 2024, and helped shape the field of agentic benchmarks by collecting together more than 100 evaluations for coding, cybersecurity, machine learning, safeguards, computer use and misalignment.
- It has been used by every major AI safety organisation
- Industry heavyweights like Weights & Biases, X.ai, Hugging Face and Alibaba have built directly on Inspect Evals to bootstrap their own evals frameworks
- Our evals are frequently used by government agencies conducting pre-deployment testing exercises
- Last year I received a notification on my phone when John Schulman raised an issue
The successful candidate will help shape the future of Inspect Evals. You’ll be embedded at the centre of the Inspect AI ecosystem, working with many highly influential and competent collaborators and users.
The main purpose of this role is to improve and drive adoption of Inspect Evals and related infrastructure.
The primary expectations for a senior product engineer are:
- Senior: self-directed, owns outcomes & delivers with autonomy; comfortable taking the lead and delegating as necessary
- Product: proactively seeks feedback and builds relationships with users; able to develop a vision and prioritise accordingly
- Engineer: able to ship high-quality software fast
In addition to core development & maintenance work, you can expect to spend your time:
- Prototyping UIs and APIs, testing them with users
- Reading research papers to understand new product requirements and use cases
- Answering questions from users in GH issues & slack messages
- Scoping work for contractors and reviewing PRs
- Building agentic workflows to streamline maintenance
- Attending and presenting at conferences and hackathons
- Providing input on work tests and hiring decisions
We encourage you to apply if this sounds like you:
- You’re passionate about scientific software, data and research
- You have 8+ years of combined experience as an engineer & scientist
- You enjoy managing projects, products and teams
- You love talking to users and helping them solve problems
- You have experience in a start-up, and thrive in fast paced, high-autonomy environments
- You’re particularly motivated by the goal of measuring the safety and risks of advanced AI
You’re exceptional fit if this sounds like you:
- You have a PhD in natural/social sciences or applied mathematics
- Your colleagues would describe you as a fast-moving, no-bullshit perfectionist
- You’re a former CTO or CEO at a scientific software start-up, where you built a team, product and user base from scratch.
- You’ve worked side-by-side with scientists, and have spent years ensuring their experiments are reliable, reproducible and easy to configure
- You’ve seen the OpenAI rogue AI hacking incidents, and you’re looking for a way to make the world safer as this technology continues to advance
- Base compensation is 110-220k GBP, with a 5k productivity budget
- We’re able to sponsor visas, and we’re open to discussing relocation packages and support
- Application deadline: 23:59 Sunday 25th October, Anywhere on Earth (AoE)
- Our team is based in London; for this role we expect at least 3 days in-person per week
Technical Staff, Cyber & Autonomous Systems Team:
- It makes it really easy to run public general capabilities evals, which we used to do in the Autonomous Systems team quite a bit and I know some other groups within AISI may be interested in doing.
- It provides a useful repository of examples of all sorts of evaluations, which makes writing new evaluations easier, especially for people getting started in Inspect.
Technical Staff, Cyber & Autonomous Systems Team:
Inspect evals is core to the work I do as a researcher. It provides access to a large repository of high-quality evals and benchmarks that are easy to run and use consistent infrastructure, allowing me to ensure experimental consistency across the work I do.
Software Engineer, Core Technology Team:
In addition to others’ points, I would add that bug reports and feature requests originating from users of Inspect Evals have been very useful to guide the development of Inspect core
Technical Staff, White Box Evaluation Team:
Inspect evals has been transformative for our research team, giving us the means to test models against a much wider range of evaluations and make stronger, more generalizable claims in our research than we would have been able to without it. The maintainers are always responsive, and the package is an excellent resource that enables the research community to run reproducible, scalable and complex evaluations with ease.
Technical Staff, Chem-Bio Team:
Inspect Evals allows me to quickly and easily test hypotheses about new models. It provides a great resource to the evals community and I’m grateful it exists.
#J-18808-Ljbffr…
