Overview
In this role you will design, deploy and maintain the infrastructure that underpins a high-performance engineering and R&D environment. You’ll work across Linux, bare-metal, cloud and containerized stacks to support compute-intensive workloads. You’ll collaborate with cross-functional teams to enable scalable, reliable HPC and AI-related compute infrastructure. The opportunity offers hands-on ownership of modern infrastructure and a chance to influence performance and reliability at scale.
Responsibilities
- Design and operate Linux-based infrastructure across bare-metal, virtualized and cloud environments
- Manage Kubernetes and containerized platforms for compute workloads
- Implement and maintain Infrastructure-as-Code using Terraform
- Configure and optimize public cloud resources (GCP or AWS)
- Automate infrastructure tasks and workflows
- Troubleshoot issues across physical and virtual layers (networking, storage, compute)
- Support HPC, GPU-enabled and research-oriented compute environments
- Collaborate with engineering/R&D teams to enable scalable compute and research workloads
Key requirements
- Strong Linux infrastructure experience
- Bare-metal server and compute environments
- Kubernetes and containerised infrastructure
- Terraform / Infrastructure-as-Code
- Public cloud infrastructure – ideally GCP or AWS
- Infrastructure automation
- Networking and storage
- CPU and/or GPU compute environments
- Troubleshooting across physical and virtual infrastructure
- Experience with HPC clusters, workload scheduling, GPU infrastructure or research/engineering compute environments
- Collaborative mindset
- Problem-solving under complex constraints
- Clear communication with engineering teams
- Linux infrastructure
- Bare-metal compute
- Kubernetes
…
