Principal Platform Engineer

Company: Trayport
Apply for the Principal Platform Engineer
Location: London
Job Description:

Overview

In this senior, hands-on role you will design, build, and operate the global platform infrastructure across cloud and on-prem environments, while acting as a technical mentor and design authority. You will lead AI-driven operations to reduce toil, speed incident response, and raise operational leverage across the Platform/Operations function. Expect to own reliable, scalable infrastructure and guide cross-team engineering with strong pragmatism and fast execution.

Responsibilities

  • Design, build, and operate AWS (VPC, networking, EKS) and Azure (AKS, AKV, networking) infrastructure, including on-prem awareness
  • Own core networking design, connectivity, DNS, load balancing, firewalls, and cloud-on-prem private links
  • Engineer for high availability with multi-region/multi-AZ architecture, capacity planning, and disaster recovery
  • Develop and maintain infrastructure-as-code (Terraform or similar), CI/CD pipelines, and Kubernetes-as-a-product
  • Define and drive SLOs, observability standards (metrics, logging, tracing) and participate in post-incident reviews
  • Continuously reduce toil through automation and scripting
  • Identify, prototype, and productionise AI-assisted workflows across the department (incident triage, runbooks, log/alert analysis)
  • Use AI-assisted engineering tools as a first-class part of your workflow and coach others to do the same safely
  • Establish guardrails for AI use in regulated, availability-critical environments

Key requirements

  • Operational/SRE background on highly available production platforms
  • Deep hands-on AWS (VPC, networking, EKS) and Azure (AKS, Key Vault, networking); multi-cloud experience
  • Strong networking fundamentals (TCP/IP, routing, firewalls, load balancing, hybrid connectivity)
  • Production Kubernetes experience at scale (day-2 operations, multi-cluster)
  • Infrastructure-as-code and automation (Terraform, Ansible; scripting in Python, Go, or Bash)
  • Demonstrable use of AI tooling to improve engineering/operational workflows
  • Credibility and communication skills to mentor senior engineers and influence design
  • Understanding of different database technologies
  • Mentorship
  • Communication
  • Influence without formal authority
  • AWS (VPC, networking, EKS)
  • Azure (AKS, Key Vault, networking)
  • Networking: TCP/IP, routing, firewalls, load balancing

…

Posted: October 1st, 2026