Senior ML Systems Engineer, Frameworks & Tooling

Company: Cohere
Apply for the Senior ML Systems Engineer, Frameworks & Tooling
Location: London
Job Description:

Overview

You will design and own the training framework for frontier-scale LLMs, blending large-scale distributed training with HPC infrastructure. The role sits at the core of ML systems, ensuring fast, reliable, and scalable model training across thousands of GPUs. You will collaborate with infra and research teams to optimize throughput and reproducibility, shaping the next generation of training infrastructure. This is a hands-on, impact-driven position in a fast-moving company solving real-world AI challenges.

Pay / Benefits

  • Open and inclusive culture
  • Health and dental benefits
  • 6 weeks vacation
  • Remote-friendly with global offices and co-working stipend
  • Parental leave top-up
  • Mental health budget and personal enrichment benefits

Responsibilities

  • Build and own the training framework for large-scale LLM training
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO, memory management, checkpointing)
  • Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100)
  • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics
  • Collaborate with infra teams to ensure Slurm setups, container environments, and hardware configurations support high-performance training
  • Investigate and resolve performance bottlenecks across the ML systems stack
  • Build robust systems that ensure reproducible, debuggable, large-scale runs

Key requirements

  • Strong engineering experience in large-scale distributed training or HPC systems
  • Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops
  • Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)
  • Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines
  • Experience working with containerized environments (Docker, Singularity/Apptainer)
  • A track record of building tools that increase developer velocity for ML teams
  • Excellent judgment around trade-offs: performance vs complexity, research velocity vs maintainability
  • Strong collaboration skills — you’ll work closely with infra, research, and deployment teams
  • Collaboration
  • Problem-solving
  • Adaptability
  • JAX internals
  • Distributed training libraries
  • Custom kernels/fused ops

…

Posted: October 1st, 2026