Senior IT Systems Administrator (Linux & HPC)

Company: KBR
Apply for the Senior IT Systems Administrator (Linux & HPC)
Location: Surrey
Job Description:

Overview

In this role you will own the Linux and HPC platform stack, delivering secure, reliable enterprise and research-computing environments. You will manage Linux systems, HPC clusters, and workload scheduling while ensuring performance and resilience at scale. You’ll work with cross-functional teams to diagnose cross-domain issues and drive continuous improvement. The role offers hands-on technical ownership and operates across on-premise and cloud-enabled HPC platforms.

Pay / Benefits

  • competitive benefits
  • professional development

Responsibilities

  • Administer, patch, harden and upgrade Linux server platforms (Red Hat Enterprise Linux or equivalent)
  • Manage core Linux services (identity, SSH, DNS, time, repos, filesystems, logging, scheduled tasks)
  • Automate tasks with shell scripts and configuration-management/orchestration tools
  • Monitor performance, availability and security; resolve cross-system issues
  • Maintain build standards, documentation and runbooks
  • Operate and support HPC clusters (SLURM: queues, partitions, policies, jobs, accounting)
  • Support NVIDIA Base Command Manager and Azure CycleCloud for cluster provisioning and lifecycle management
  • Collaborate with engineers and users to diagnose job, compiler and MPI issues; plan maintenance activities
  • Administer hardware lifecycle, Cisco compute platforms, and NetApp storage integration
  • Troubleshoot networking, storage connectivity and end-to-end dependencies
  • Apply secure configuration, backup/DR recovery, incident response and change management processes
  • Provide technical guidance and knowledge transfer to colleagues and service-desk teams

Key requirements

  • Hands-on Linux administration experience in complex enterprise or research environments
  • Experience supporting HPC clusters and diagnosing cross-layer issues
  • Strong SLURM administration and workload troubleshooting
  • Bash or similar scripting for task automation
  • Experience with Linux performance, patching, security hardening and vulnerability remediation
  • Experience with shared file services (NFS, permissions, throughput)
  • Networking fundamentals and distributed-system dependencies
  • Experience delivering controlled technical change and maintaining documentation in ITSM
  • Bachelor’s degree in computing, engineering or related field or equivalent
  • Desirable: experience with NVIDIA Base Command Manager, Bright Cluster Manager, NetApp, InfiniBand, MPI workloads, Ansible, Git, virtualization/containers
  • Desirable: certifications in Linux, HPC, Cisco, NVIDIA (advantageous)
  • Independent working style
  • Strong communication to specialists and non-specialists
  • Attention to detail and operational discipline
  • Linux administration (RHEL or equivalent)
  • HPC cluster management
  • SLURM workload scheduler

…

Posted: September 20th, 2026