Overview
You will own and build the AI training infrastructure at a pioneering robotics AI startup. You’ll design and scale platforms that enable researchers to translate ideas into production-ready systems. The role sits at the intersection of machine learning, distributed systems, cloud computing, and robotics, with the goal of accelerating breakthroughs in real-world robotics. You’ll join a founding, highly collaborative team and influence the technical direction from an early stage.
Responsibilities
- Design and scale systems for large-scale model training
- Improve efficiency of distributed workloads
- Develop tooling to enable researchers to experiment rapidly and deploy ideas
- Build a platform that may be novel in the field
- Own infrastructure powering AI training and production readiness
- Work across research and engineering to solve complex problems and shape technical direction
Key requirements
- Experience building high-performance infrastructure for machine learning or compute-intensive workloads
- Experience with distributed training, GPU orchestration, cloud infrastructure
- Knowledge of modern deep learning frameworks and production software engineering
- Background in AI infrastructure, platform engineering, or related fields (valued)
- Ability to solve complex technical problems over focusing on titles
- Curiosity
- Strong technical depth
- Motivation to tackle difficult problems
- Distributed training
- GPU orchestration
- Cloud infrastructure
…
