Director: AI Systems Reliability, Testing & Performance

Company: Novartis
Apply for the Director: AI Systems Reliability, Testing & Performance
Location: London
Job Description:

Overview

In this senior AI leadership role, you’ll shape how AI systems across Novartis Development are evaluated, tested, and validated, driving evidence of reliability, security, and fitness-for-use. You’ll lead evaluation, benchmarking, testing, and production-readiness across diverse AI capabilities, from models to digital twins. Partnering with engineering and cybersecurity teams, you’ll establish observability, drift detection, and performance intelligence to support regulated, business-critical AI. You’ll champion rigorous, transparent methods that enable confident deployment and continuous improvement.

Pay / Benefits

  • flexible and hybrid working options
  • minimum 14 weeks paid parental leave
  • insurance plans
  • retirement plans
  • wellbeing resources
  • global recognition programs

Responsibilities

  • Define evaluation methodologies for AI systems, models, agents, digital twins, and simulation environments across Development
  • Create benchmark suites, datasets, and testing harnesses for AI initiatives
  • Set objective measures for quality, reliability, robustness, grounding, and task success
  • Ensure evaluation approaches are rigorous, reproducible, and comparable across AI systems
  • Build a common evidence framework for AI capabilities, limitations, and fitness-for-use
  • Define debugging and validation approaches for agentic systems, digital twins, and human-in-the-loop workflows
  • Develop production readiness criteria for AI in Development environments
  • Support validation activities required for GxP-relevant and regulated AI systems
  • Ensure AI systems are tested under realistic operating conditions with objective validation evidence
  • Define how reliability, drift, and performance are measured and monitored over time
  • Establish monitoring and observability requirements across Development AI
  • Collaborate with engineering to provide required signals and traceability
  • Identify and diagnose AI system failures and promote evidence-based improvement
  • Define AI security testing approaches including red teaming and adversarial testing
  • Assess AI vulnerabilities with cybersecurity teams and include security in validation
  • Create portfolio-wide standards for capturing evidence on AI quality, reliability, and performance
  • Define standard reporting for benchmark results, validation findings, monitoring signals, and reliability assessments
  • Provide transparency into AI strengths, limitations, and failure modes
  • Enable objective decisions on readiness and fitness-for-use

Key requirements

  • 10+ years in AI, ML, software engineering, quality, risk, validation, or technology governance
  • Experience deploying or overseeing business-critical AI systems
  • Experience establishing governance, controls, quality standards, or operational frameworks
  • Experience defining evaluation approaches and performance metrics
  • Experience managing operational risk in regulated environments
  • Experience partnering with infrastructure and platform teams
  • Stakeholder engagement
  • Mentorship
  • Change management
  • AI evaluation and benchmarking
  • AI testing and validation
  • Reliability engineering and observability

Posted: September 14th, 2026