Senior Linux Administrator

Company: Riverlane
Apply for the Senior Linux Administrator
Location: Cambridge
Job Description:

Overview

In this role you will own and grow Riverlane’s core HPC and Linux infrastructure to support cutting-edge quantum error correction research. You will lead day-to-day HPC cluster administration while shaping long-term plans for a scalable, fault-tolerant environment. You’ll work closely with the Infrastructure Team to ensure reliability, performance, and security for high-demand compute workloads. This is a chance to impact foundational software and hardware integration in a fast-paced, mission-driven company.

Pay / Benefits

  • annual bonus plan
  • private medical insurance
  • life insurance
  • contributory pension scheme
  • equity
  • 28 days annual leave

Responsibilities

  • Administer and maintain the HPC cluster (compute nodes, storage, networking)
  • Evolve cluster topology, node configuration and resource utilisation as the estate grows
  • Deploy and troubleshoot HPC tooling, especially Slurm (queues and scheduling policies)
  • Monitor performance and ensure high availability using observability tools (Prometheus)
  • Manage network file storage for throughput and latency
  • Patch, support and troubleshoot core Linux infrastructure (RHEL)
  • Escalation point for complex issues; support HPC users
  • Maintain Synopsys EDA tooling including licence management and performance tuning
  • Improve self-service capabilities for users (e.g., password resets, VNC sessions)
  • Automate routine tasks with Bash, Python, Ansible; plan improvements with Infrastructure Team
  • Own work packages, meet sprint goals; build and maintain system documentation
  • Ensure security, hardening, lifecycle management and disaster recovery readiness
  • Maintain off-site backup strategy and data retention compliance
  • Provide stakeholder updates on progress and improvements

Key requirements

  • Extensive experience administering infrastructure in fast-paced engineering environments
  • Enterprise Linux expertise (RHEL) including hardening and lifecycle management
  • Hands-on with Slurm or similar job schedulers in distributed compute environments
  • Strong understanding of network file storage and workload optimization
  • Proficiency with containers and virtualization infrastructure
  • Automation using Ansible, Python, Bash; infrastructure-as-code practices
  • Experience with backup/recovery strategies and DR planning
  • Experience managing licensed EDA tools (e.g., Synopsys) is a strong plus
  • Excellent communication skills for technical and non-technical users
  • strong communication
  • supportive with users
  • team collaboration
  • Slurm scheduling
  • RHEL administration and hardening
  • HPC cluster architecture

…

Posted: October 10th, 2026