Overview
Join McKinsey’s data engineering team in London to build scalable data foundations for cutting-edge AI systems. You will partner with cross-functional units to deliver production-ready data pipelines and secure data environments that power enterprise AI. The role blends hands-on engineering with R&D to scale agentic and generative AI capabilities for client impact. You’ll learn rapidly in a high-performance culture and contribute to shaping AI solutions at scale.
Pay / Benefits
- competitive salary
- comprehensive benefits package
- global exposure
- structured learning and apprenticeship culture
- mentorship and career development
- inclusive, diverse workforce
Responsibilities
- Build foundational data infrastructure powering AI applications (LLMs, retrieval systems, workflows)
- Design and maintain scalable data pipelines and secure data environments
- Prepare data for AI-driven systems and collaborate with cross-functional teams
- Develop scalable, reproducible data components for ML, agentic, and autonomous AI
- Assess data landscapes and data quality; translate hypotheses into engineered features
- Contribute to R&D initiatives to innovate and scale AI capabilities
- Collaborate with McKinsey QuantumBlack, AI by McKinsey, and QuantumBlack Labs
- Support client-facing technologist work and cross-functional agile delivery
Key requirements
- Degree in Computer Science/Engineering, or equivalent experience
- Experience in a data-focused role (internships, academic projects)
- Proficiency in Python and SQL
- Exposure to Agentic AI, Generative AI, ML, or BI across data formats and processing methods
- Familiarity with data platforms (Databricks, Snowflake, BigQuery, PSQL, etc.) and cloud platforms (AWS, Azure, GCP)
- Experience with Pandas, Spark, dbt, LangChain, etc.
- Knowledge of Git, DevOps and MLOps/LLMOps concepts, CI/CD
- Strong verbal and written communication in English
- Willingness to learn quickly and adapt to different tech stacks
- Experience with coding agents (Cursor, Claude Code, Codex) is a plus
- Strong communication
- Time management in autonomous environments
- Adaptability and fast learning
- Python
- SQL
- Pandas
…
