Founding GPU Engineer

Company: Fuse Energy Supply
Apply for the Founding GPU Engineer
Location: London
Job Description:

Overview

As Founding GPU Engineer, you will architect and optimize CUDA-based software for data center AI workloads. You’ll connect GPU performance with energy availability and grid signals, enabling efficient, scalable compute at scale. You’ll work across low-level kernels to system-level infrastructure with cross-functional teams. This role offers a chance to shape a mission-driven energy and AI platform with real-world impact. You help drive performance and energy efficiency at the physics of data centers.

Pay / Benefits

  • competitive salary and an equity sign-on bonus
  • Biannual bonus scheme
  • Fully expensed tech
  • Breakfast and dinner allowance

Responsibilities

  • Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads
  • Profile and tune GPU performance across compute, memory bandwidth, and interconnect bottlenecks
  • Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals
  • Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar libraries
  • Collaborate with data center infrastructure teams on power capping and dynamic voltage/frequency scaling
  • Integrate custom kernels into training/inference pipelines with ML/systems engineers
  • Benchmark against CPU/GPU baselines and drive continuous performance improvements
  • Contribute to internal libraries, documentation, and best practices for GPU performance engineering

Key requirements

  • 4+ years of experience writing production CUDA code
  • Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy)
  • Proficiency in C++ and CUDA; Python for tooling
  • Experience with performance profiling tools (Nsight Systems/Compute)
  • Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand)
  • Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design
  • Comfortable working across the stack from low-level kernels to system-level infrastructure
  • CUDA
  • C++
  • Python
  • Nsight Systems/Compute
  • NCCL
  • MPI

…

Posted: October 1st, 2026