Overview
In this senior AI leadership role, you’ll shape how AI systems across Novartis Development are evaluated, tested, and validated, driving evidence of reliability, security, and fitness-for-use. You’ll lead evaluation, benchmarking, testing, and production-readiness across diverse AI capabilities, from models to digital twins. Partnering with engineering and cybersecurity teams, you’ll establish observability, drift detection, and performance intelligence to support regulated, business-critical AI. You’ll champion rigorous, transparent methods that enable confident deployment and continuous improvement.
Pay / Benefits
- flexible and hybrid working options
- minimum 14 weeks paid parental leave
- insurance plans
- retirement plans
- wellbeing resources
- global recognition programs
Responsibilities
- Define evaluation methodologies for AI systems, models, agents, digital twins, and simulation environments across Development
- Create benchmark suites, datasets, and testing harnesses for AI initiatives
- Set objective measures for quality, reliability, robustness, grounding, and task success
- Ensure evaluation approaches are rigorous, reproducible, and comparable across AI systems
- Build a common evidence framework for AI capabilities, limitations, and fitness-for-use
- Define debugging and validation approaches for agentic systems, digital twins, and human-in-the-loop workflows
- Develop production readiness criteria for AI in Development environments
- Support validation activities required for GxP-relevant and regulated AI systems
- Ensure AI systems are tested under realistic operating conditions with objective validation evidence
- Define how reliability, drift, and performance are measured and monitored over time
- Establish monitoring and observability requirements across Development AI
- Collaborate with engineering to provide required signals and traceability
- Identify and diagnose AI system failures and promote evidence-based improvement
- Define AI security testing approaches including red teaming and adversarial testing
- Assess AI vulnerabilities with cybersecurity teams and include security in validation
- Create portfolio-wide standards for capturing evidence on AI quality, reliability, and performance
- Define standard reporting for benchmark results, validation findings, monitoring signals, and reliability assessments
- Provide transparency into AI strengths, limitations, and failure modes
- Enable objective decisions on readiness and fitness-for-use
Key requirements
- 10+ years in AI, ML, software engineering, quality, risk, validation, or technology governance
- Experience deploying or overseeing business-critical AI systems
- Experience establishing governance, controls, quality standards, or operational frameworks
- Experience defining evaluation approaches and performance metrics
- Experience managing operational risk in regulated environments
- Experience partnering with infrastructure and platform teams
- Stakeholder engagement
- Mentorship
- Change management
- AI evaluation and benchmarking
- AI testing and validation
- Reliability engineering and observability
…
