Overview
In this role you will design and own the large-scale AI training infrastructure powering robotics-focused models. You will scale distributed training across GPUs and clouds, optimize throughput and costs, and turn vast simulation data into production-ready models. You’ll build tooling to accelerate research, collaborating closely with founders and the research team to tackle frontier challenges. This is a chance to shape the architecture and technical direction of a high-impact robotics AI startup from day one.
Pay / Benefits
- Meaningful equity
- Founding-team influence
- Collaborative, small team environment
Responsibilities
- Design and own large-scale AI training infrastructure
- Scale distributed model training across hundreds of GPUs and multiple cloud environments
- Optimize training throughput, GPU utilization and infrastructure costs
- Build systems translating simulation data into production-ready models
- Develop tooling and frameworks to accelerate research and experimentation
- Collaborate with founders and research team to resolve bottlenecks at the frontier of robotics
- Shape the architecture, culture and technical direction of the company from day one
Key requirements
- Large-scale distributed training experience
- PyTorch and modern deep learning frameworks
- Kubernetes, Slurm or GPU orchestration platforms
- AWS and specialist GPU cloud providers
- High-performance computing and distributed systems
- Training optimisation, memory management and networking
- MLOps tooling, CI/CD and experimentation frameworks
- Python and production-grade software engineering
- Computer vision, multimodal AI or transformer architectures
- Deeply technical and problem-solving mindset
- Thrives in small, ambitious teams
- Enjoys building from scratch and avoiding legacy systems
- Large-scale distributed training
- PyTorch and modern DL frameworks
- Kubernetes/Slurm or GPU orchestration
…
