Agentic AI, Data Engineer COA Accelerator

Company: IQVIA
Apply for the Agentic AI, Data Engineer COA Accelerator
Location: London
Job Description:

Overview

In this role you will design and maintain scalable data infrastructure to power IQVIA’s AI-enabled COA strategy. You will own end-to-end data ingestion, transformation, enrichment, indexing and governance across public, proprietary and client sources. You will build pipelines for structured and unstructured content (PDFs, documents, registries) turning raw material into AI-ready, retrievable assets, with strong quality controls and traceability. You will collaborate with AI engineers and cross-functional teams to deliver reliable, compliant knowledge platforms that support evidence generation.

Pay / Benefits

  • remote or hybrid work
  • career development
  • collaborative, multi-cultural environment

Responsibilities

  • Design, build, and maintain data infrastructure for AI-enabled COA workflows
  • Own ingestion, transformation, normalisation, enrichment, indexing, versioning, and governance of data sources
  • Create ingestion pipelines for structured and unstructured sources (PDFs, slides, docs, databases, APIs, registries, publications)
  • Transform material into standardised, AI-ready formats for retrieval, citations, and expert review
  • Develop repeatable parsing, OCR, metadata enrichment, chunking, deduplication, versioning, and quality control processes
  • Support core knowledge layer data (instrument metadata, translations, usage rights, psychometric data, mappings, regulatory precedent)
  • Assist integration of public and proprietary data sources for evidence repositories
  • Prepare data for retrieval-augmented generation with high-quality chunking, embeddings, indexes, and filters
  • Collaborate with AI engineers to improve retrieval precision, relevance, and citation accuracy
  • Maintain traceability and data quality controls; keep audit trails for ingestion, transformations, and access
  • Work with legal, security, compliance, product, and domain stakeholders to ensure governance and licensing compliance
  • Collaborate across AI Engineering, Product, COA Science and Software Engineering

Key requirements

  • Degree in computer science, data engineering, data science, information systems, bioinformatics, computational biology, engineering, or related field
  • Experience building data pipelines for structured and unstructured data
  • Strong Python and SQL skills
  • Experience with APIs, relational databases, document stores, search indexes, cloud data platforms, and ETL/ELT
  • Experience handling large volumes of text-heavy documents (PDFs, slides, reports, publications, regulatory files)
  • Knowledge of data cleaning, normalisation, metadata management, document parsing, indexing, lineage, versioning, and auditability
  • Familiarity with data modelling for scientific, clinical, regulatory or healthcare content
  • Ability to translate domain requirements into practical data structures and reusable pipelines
  • Strong attention to detail and ability to identify data quality issues affecting AI outputs and traceability
  • Ability to collaborate with diverse teams and document data sources, transformations, and rules
  • collaborative mindset
  • strong communication
  • attention to detail
  • Python
  • SQL
  • APIs

…

Posted: October 1st, 2026