HPC Network Engineer

Company: Fuse Energy Supply
Apply for the HPC Network Engineer
Location: London
Job Description:

Overview

You design, deploy, and operate the network fabric for a multi-tenant AI cluster, enabling high-performance GPU compute and storage. You own both data center and office networks, shaping architecture through day-2 operations and driving reliability at scale. You collaborate with cross-functional teams to ensure secure, scalable, and observable fabric delivery. This role combines hands-on engineering with documentation, mentoring, and continuous improvement to support Fuse Energy’s mission.

Pay / Benefits

  • Competitive salary and an equity sign-on bonus
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Breakfast and dinner allowance for office-based employees

Responsibilities

  • Design and operate lossless, RDMA-capable fabrics (RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS and buffer tuning at scale
  • Build and manage leaf-spine data center fabrics with routed underlay and overlay design (BGP, EVPN/VXLAN)
  • Implement and maintain per-tenant network isolation across compute, storage, and management planes
  • Automate network provisioning, configuration, and validation using code (Ansible, Python, NetBox) deployed via CI
  • Develop telemetry and observability for the fabric with dashboards and alerts (Prometheus/Grafana/Datadog, streaming telemetry)
  • Troubleshoot end-to-end performance from optics to NIC/DPU configuration and host behavior
  • Operate the out-of-band management network, console access, and remote recovery paths
  • Support tenant onboarding: segmentation, bandwidth guarantees, and capacity planning as the cluster scales
  • Write clear design documentation detailing decisions and rationale
  • Own and maintain the office network: wired/wireless, firewalling, VPN/remote access
  • Upskill colleagues through documentation and hands-on sessions to improve fabric operation

Key requirements

  • 5+ years operating production data center networks
  • Strong dynamic routing experience (BGP) and overlay design (EVPN/VXLAN)
  • Hands-on experience with leaf-spine/Clos fabric design and operation
  • Proficiency with modern data center network operating systems and Linux networking stack
  • Practical RDMA fabric experience (RoCEv2 or InfiniBand) and understanding of lossless behavior for GPU workloads
  • Network automation experience (Python, Ansible), config management and source-of-truth practices
  • Solid Linux administration fundamentals
  • Experience with network telemetry and monitoring (Prometheus/Grafana, sFlow/IPFIX, streaming telemetry)
  • Experience with corporate/campus networks (wired/wireless, NAC/802.1X, VPN)
  • Clear communicator who enjoys teaching with documentation, pairing, and knowledge sharing
  • Clear communicator and teacher
  • Collaborative mindset and ability to pair with less experienced colleagues
  • Ability to document, share knowledge, and run training sessions
  • RDMA fabrics (RoCEv2, InfiniBand)
  • lossless Ethernet concepts (PFC/ECN/DCQCN)
  • leaf-spine/Clos fabric design

…

Posted: September 14th, 2026