Lead AI Infrastructure & Distributed Systems Engineer

Company: LinuxRecruit
Apply for the Lead AI Infrastructure & Distributed Systems Engineer
Location: London
Job Description:

Overview

In this early-stage robotics startup, you will own the core model training infrastructure, building scalable systems to train foundation-model style AI for physical tasks. You’ll optimize GPU clusters, orchestration, and data pipelines to remove bottlenecks and accelerate research. This hands-on role sits at the crossroads of HPC, MLOps, and high-performance networking, with full architectural ownership from day one. You’ll work in person in central London as part of a small elite team, with substantial equity upside as the company scales. You drive impact by turning ambitious research into production-ready infrastructure.

Pay / Benefits

  • equity upside
  • in-person role in central London
  • founding-level team
  • small elite group
  • startup environment

Responsibilities

  • Design scalable distributed training infrastructure from scratch
  • Eliminate bottlenecks in training pipelines and accelerate research cycles
  • Optimize GPU compute, cluster scheduling, and high-performance networking
  • Maintain and improve MLOps and AI infrastructure practices
  • Work hands-on with AWS, Kubernetes, Slurm, and PyTorch in production environments
  • Own architectural decisions and drive system reliability and performance

Key requirements

  • Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks
  • Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking
  • Experience in MLOps, AI infrastructure, or HPC
  • Prior robotics experience not required; CV pipelines are advantageous
  • In-person, founding-level collaborator in central London
  • Equity upside potential
  • high agency
  • ownership and accountability
  • fast-paced startup mindset
  • AWS
  • Kubernetes
  • Slurm

…

Posted: October 1st, 2026