Overview
As Platform Engineer in the UK Digital Competence Centre, you will support an on-premise AI platform, ensuring security, reliability and performance. You’ll work within an agile platform team to maintain the core infrastructure and enable cross-functional delivery with engineers, security and service teams. The role blends operations, lifecycle management and continuous improvement to keep critical hosted solutions available. You will contribute to documentation and standards while helping scale and govern the platform. This is a hands-on, impact-focused opportunity in a mission-driven, security-conscious environment.
Pay / Benefits
- private medical insurance
- buying or selling annual leave
- cycle to work schemes
- employee discounts
- paid volunteering day
- stocks and shares
Responsibilities
- Support and maintain on-premise hosted AI platform infrastructure to ensure security, reliability and availability
- Monitor platform health, investigate and resolve incidents as part of day-to-day support
- Perform routine maintenance and lifecycle activities (patching, upgrades, scale-out)
- Collaborate with engineering, security and service teams to ensure resilient, well-governed platform
- Contribute to technical documentation and operational standards
- Operate day-to-day platform across physical servers, virtualization, and OS environments
- Support containerised services and underlying infrastructure
- Execute maintenance activities including patching, upgrades and health checks
- Manage hardware/software lifecycles to maintain security and supportability
- Maintain technical documentation and operational records
- Investigate and resolve incidents per support processes; work with teams and suppliers to address recurring problems
- Identify opportunities to improve platform reliability, supportability and efficiency
Key requirements
- Experience with Windows Server operating systems
- Experience with Unix-based operating systems, especially Red Hat Enterprise Linux (RHEL)
- Experience with VMware ESXi and/or other hypervisors
- Experience supporting containerised environments, especially Kubernetes
- Incident and problem management knowledge; ability to follow change management processes
- Good understanding of core networking principles
- Understanding of application and OS lifecycle management
- Experience supporting infrastructure in an enterprise or operational environment
- Good understanding of platform and infrastructure support principles
- BitBucket/Github CI/CD tools
- Terraform or Ansible (Infrastructure as Code)
- Scripting languages (desirable)
- HP Enterprise hardware
- Agile methodologies
- Active Directory
…
