Machine Learning Performance Engineer

Company: G Research
Apply for the Machine Learning Performance Engineer
Location: London
Job Description:

Overview

In this role you optimize large-scale ML workloads on GPU and CPU infrastructure to accelerate researchers’ work. You’ll profile, benchmark, and refine distributed workloads, shaping tools and libraries that improve efficiency. You collaborate with researchers and platform teams to guide long-term platform evolution, balancing performance and reliability. This is a hands-on opportunity to impact scale, architecture, and tooling for cutting-edge ML computation at a research-driven firm.

Pay / Benefits

  • competitive compensation with discretionary bonus
  • lunch provided
  • barista bar
  • 35 days annual leave
  • 9% pension contributions
  • healthcare and life assurance

Responsibilities

  • Profile, benchmark and tune large-scale training and inference workloads across CPU, GPU and memory-intensive systems
  • Collaborate with researchers and engineers to design optimized compute solutions
  • Develop reference implementations, libraries and tools to improve job efficiency and reliability
  • Work with systems, architecture and platform teams to evolve the compute stack
  • Influence long-term platform and infrastructure decisions

Key requirements

  • Bachelor/Master/PhD in computer science or equivalent experience
  • Proven track record of profiling, benchmarking and optimizing distributed workloads
  • Experience with Python
  • Knowledge of CUDA
  • Experience with HPC schedulers and Kubernetes-based workload orchestration
  • Strong understanding of PyTorch or other deep learning frameworks
  • Strong background in data structures, algorithms, and parallel programming on heterogeneous systems
  • Deep understanding of Linux fundamentals (scheduling, memory, NUMA, networking, filesystems)
  • Familiarity with profiling/monitoring tools (nsys, ncu, eBPF-based tools, performance counters)
  • Strong communication and cross-team collaboration skills
  • collaboration
  • clear communication
  • problem-solving
  • Python
  • CUDA
  • HPC schedulers

…

Posted: October 1st, 2026