CUDA Engineer

Company: Fuse Energy Supply
Apply for the CUDA Engineer
Location: London
Job Description:

Overview

As a CUDA Engineer at Fuse Energy, you will design and optimize low-level GPU kernels that power transformer inference workloads in a fast-growing renewable energy and AI compute environment. You will push throughput and latency, tune memory hierarchies, and implement quantisation-aware, mixed-precision kernels across cutting-edge GPUs. Your work directly supports energy-efficient AI inference at scale, helping Fuse meet rapid compute demands while aligning with the company’s mission to accelerate clean energy via advanced technology. You will collaborate with cross-functional teams to improve performance, reliability, and efficiency of our inference pipelines.

Pay / Benefits

  • Competitive salary and an equity sign-on bonus
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Breakfast and dinner allowance for office based employees

Responsibilities

  • Write and optimize custom CUDA kernels for core transformer inference operations
  • Profile kernels to identify bottlenecks in occupancy, memory throughput, and warp divergence
  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines
  • Optimize memory access patterns and manage the memory hierarchy for maximum bandwidth utilization
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint
  • Build and tune caching mechanisms for efficient autoregressive decoding
  • Tune kernel launch configurations for target GPU architectures
  • Benchmark kernels against baselines and drive throughput/latency improvements
  • Write tests for CUDA code to catch performance and correctness regressions
  • Maintain internal CUDA libraries and contribute to team coding standards and documentation

Key requirements

  • 4+ years writing production CUDA code with a track record of shipping performance-critical kernels
  • Deep understanding of GPU microarchitecture, warps, occupancy, register pressure, and memory hierarchy
  • Strong CUDA C++ skills, including streams and asynchronous execution
  • Hands-on experience profiling to diagnose compute-bound vs memory-bound bottlenecks
  • Experience with kernel fusion, memory coalescing, and avoiding warp divergence
  • Experience writing quantised and mixed-precision kernels
  • Solid grasp of parallel algorithm design and numerical precision tradeoffs
  • CUDA C++
  • GPU profiling
  • kernel fusion
  • memory hierarchy optimization
  • quantisation-aware kernels
  • mixed-precision arithmetic

…

Posted: October 1st, 2026