AI Inference Engineer

Company: Fuse Energy Supply
Apply for the AI Inference Engineer
Location: London
Job Description:

Overview

In this Founding role, you will define and build Fuse’s AI inference serving layer to scale models with high throughput and low latency. You’ll own the architecture and strategy for serving, collaborating with CUDA/GPU engineers to integrate performance improvements. You’ll translate targets into concrete capacity plans and reliability standards, shaping a critical function at the intersection of energy and AI. This opportunity offers the chance to influence a foundational platform at a fast-growing renewable energy startup and work directly with the CTO to drive impact.

Pay / Benefits

  • Competitive salary and equity sign-on bonus
  • Biannual bonus scheme
  • Fully expensed tech
  • Breakfast and dinner allowance

Responsibilities

  • Define inference serving strategy and architecture from first principles
  • Design and build the serving stack: routing, batching, scheduling, autoscaling for high-throughput, latency-sensitive workloads
  • Own model-level optimisation strategy (quantisation, distillation, speculative decoding) in collaboration with CUDA/GPU teams
  • Make core architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents)
  • Translate throughput, latency, and uptime commitments into technical specs and capacity plans
  • Act as the technical owner of inference performance and reliability
  • Work closely with CUDA/GPU teams to integrate low-level performance work into the serving layer
  • Set standards, tooling, and benchmarks for the function’s growth

Key requirements

  • 4+ years building or operating large-scale inference serving systems or equivalent experience
  • Deep, hands-on experience with inference serving frameworks and optimization techniques (batching, KV-cache management, quantisation, speculative decoding)
  • Strong systems thinking across a large cluster from request to response
  • Experience collaborating with GPU/CUDA engineers to integrate performance work into serving systems
  • Track record of making high-stakes architecture decisions and owning outcomes
  • Comfort operating without a playbook in an early-stage, founding role
  • strong communication
  • ownership
  • collaborative problem-solving
  • inference serving frameworks
  • throughput and latency optimization
  • quantisation

…

Posted: October 1st, 2026