Overview
As Senior AI Platform Engineer, you define and execute the infrastructure strategy behind IQVIA’s LLM programmes, delivering secure, scalable, production-ready AI solutions. You’ll provide technical leadership across compute, data, model lifecycle, and platform engineering, partnering with research, product, data engineering, and MLOps teams to support training, evaluation, deployment, and governance of large-scale AI systems. This role offers impact through enabling rapid experimentation with operational excellence, governance, and reliability. You’ll work in a cross-functional environment shaping platform choices and collaborations that advance healthcare-focused AI initiatives.
Pay / Benefits
- access to cutting-edge technology
- diverse geographies and teams
- clear career development path
- opportunity to impact patient care
- growth and learning culture
- global company with healthcare focus
Responsibilities
- Own the AI platform and infrastructure roadmap for LLM initiatives
- Design and deliver high-performance compute environments across AWS and on-premises (GPU, Slurm)
- Optimize training and inference workloads for performance, scalability, and reliability
- Establish model and data lifecycle capabilities (versioning, lineage, reproducibility, registries)
- Evolve knowledge graph infrastructure and integrate with AI workflows
- Coordinate technically across AI Research, Data Engineering, MLOps, Product, and Infrastructure
- Lead vendor selection, procurement, and technology partnerships for compute architectures and AI platforms
- Define platform engineering standards, governance, and best practices; mentor engineers
Key requirements
- Experience designing, building, and operating large-scale AI/distributed computing platforms in enterprise
- Deep understanding of LLM architectures and GPU tooling (CUDA, cuDNN, NCCL) and distributed training (PyTorch)
- Knowledge of distributed training/inference strategies (tensor, pipeline, data, expert parallelism)
- Experience optimizing LLM inference with vLLM, TensorRT-LLM, NVIDIA NIM, SGLang or similar
- Model optimization techniques (quantisation, mixed precision FP8, LoRA) and deployment tuning
- GPU workload profiling/troubleshooting (Nsight, DCGM)
- Strong AWS background; HPC, distributed systems, containers, automation
- Experience with Slurm, Kubernetes, Ray or similar orchestration
- Proven ability to bridge research and production with governance, security, reliability
- Effective communication with engineers, researchers, product and executives
- Cross-functional collaboration
- Strategic thinking and stakeholder management
- Technical leadership and mentoring
- LLM architectures and GPU integration
- CUDA, cuDNN, NCCL
- PyTorch and distributed training
…
