AI Infrastructure Engineer

Company: Intercom
Apply for the AI Infrastructure Engineer
Location: London
Job Description:

Overview

In this role you will help build and scale Intercom’s AI infrastructure, enabling fast training and reliable inference for large transformer and LLM models. You will work with a small, highly technical team to push the performance and reliability of our AI platform, powering Fin and the Intercom Customer Service Suite. You’ll collaborate with ML scientists to bring cutting-edge training and inference methods into production and mentor other engineers. This is a hands-on, impact-driven role at the core of our AI strategy.

Pay / Benefits

  • Competitive salary and equity
  • Lunch and snacks
  • Performance reviews
  • Access to Claude Code and AI tools
  • Pension scheme with 4% match
  • Health and dental insurance for you and dependents

Responsibilities

  • Implement and scale training pipelines for large transformer and LLM models from data ingestion to distributed training and evaluation
  • Build and optimize low-latency, reliable inference services with autoscaling, routing, and fallbacks
  • Tune GPU kernels, improve utilization, and identify bottlenecks across training and inference stacks
  • Collaborate with ML scientists to productionize advanced training and inference methods
  • Mentor and develop other engineers on the team and elevate technical standards and reliability across the AI platform

Key requirements

  • 5+ years of software engineering with a track record of shipping high-quality products or platforms
  • Degree in Computer Science, Computer Engineering, or related field (or equivalent experience)
  • Hands-on experience with model training (transformers/LLMs) or model inference at scale
  • Experience with low-level GPU work (CUDA, Triton)
  • Production experience at meaningful scale and strong communication skills
  • Proficiency with at least one programming language (Python, Ruby, Java, Go, etc.)
  • Clear communication
  • Collaborative mindset
  • Continuous learner
  • Model training with transformers/LLMs
  • Model inference at scale
  • GPU programming (CUDA, Triton)

…

Posted: September 30th, 2026