Overview
Join a stealth AI robotics startup as part of the AI infrastructure team and own the systems powering large-scale model training. You will architect and optimize distributed workloads across multi-node GPU clusters in AWS, tackling bottlenecks to maximize throughput and cost-effectiveness. You’ll manage cluster orchestration with Slurm and Kubernetes while evolving the platform for next-generation GPU infrastructure. Work across the full AI training stack with a world-class team in the London lab, shaping the company’s future breakthroughs. This is a high-impact, early-stage opportunity with significant ownership and influence.
Responsibilities
- Own systems powering large-scale model training
- Architect and optimize distributed workloads across multi-node GPU clusters in AWS
- Remove bottlenecks in the training pipeline to maximize GPU utilization and throughput
- Improve cost-efficiency of training workflows
- Manage cluster orchestration with Slurm and Kubernetes
- Evolve platform to support next-generation GPU infrastructure and specialized compute providers
- Work across the full AI training stack from PyTorch-based frameworks to networking and storage
- Collaborate with the London lab to influence core system design and technical direction
Key requirements
- AWS
- Slurm
- Kubernetes
- PyTorch
- distributed training
- high-performance infrastructure
…
