Overview
In this role you will help build and scale Intercom’s AI infrastructure, enabling fast training and reliable inference for large transformer and LLM models. You will work with a small, highly technical team to push the performance and reliability of our AI platform, powering Fin and the Intercom Customer Service Suite. You’ll collaborate with ML scientists to bring cutting-edge training and inference methods into production and mentor other engineers. This is a hands-on, impact-driven role at the core of our AI strategy.
Pay / Benefits
- Competitive salary and equity
- Lunch and snacks
- Performance reviews
- Access to Claude Code and AI tools
- Pension scheme with 4% match
- Health and dental insurance for you and dependents
Responsibilities
- Implement and scale training pipelines for large transformer and LLM models from data ingestion to distributed training and evaluation
- Build and optimize low-latency, reliable inference services with autoscaling, routing, and fallbacks
- Tune GPU kernels, improve utilization, and identify bottlenecks across training and inference stacks
- Collaborate with ML scientists to productionize advanced training and inference methods
- Mentor and develop other engineers on the team and elevate technical standards and reliability across the AI platform
Key requirements
- 5+ years of software engineering with a track record of shipping high-quality products or platforms
- Degree in Computer Science, Computer Engineering, or related field (or equivalent experience)
- Hands-on experience with model training (transformers/LLMs) or model inference at scale
- Experience with low-level GPU work (CUDA, Triton)
- Production experience at meaningful scale and strong communication skills
- Proficiency with at least one programming language (Python, Ruby, Java, Go, etc.)
- Clear communication
- Collaborative mindset
- Continuous learner
- Model training with transformers/LLMs
- Model inference at scale
- GPU programming (CUDA, Triton)
…
