Production Software Engineer

Company: G Research
Apply for the Production Software Engineer
Location: London
Job Description:

Overview

In this role you will drive resilience, observability and runtime efficiency for a real-time distributed platform. You will design tooling and workflows to support fast, safe software delivery and deploy automation across a federated engineering environment. You’ll work with software, infrastructure, front-office and research teams to improve stability, reduce operational risk and enable continuous improvement. This is a high-impact, hands-on role with a clear focus on production reliability and performance.

Pay / Benefits

  • Highly competitive compensation plus annual discretionary bonus
  • Lunch provided (Just Eat for Business) and barista bar
  • 35 days’ annual leave
  • 9% company pension contributions
  • Informal dress code and good work-life balance
  • Comprehensive healthcare and life assurance cycle-to-work scheme

Responsibilities

  • Improve resilience and efficiency of real-time distributed systems by reducing bottlenecks and toil
  • Develop tooling and frameworks for frequent, low-risk software delivery
  • Own metrics, alerting and diagnostics infrastructure for monitoring the platform
  • Build and maintain deployment automation, observability and runtime management systems
  • Promote runtime engineering best practices across federated teams and establish fault tolerance standards
  • Participate in shared production support rotation to respond to incidents and drive improvements
  • Collaborate with application, research and execution teams to define runtime boundaries and production SLAs

Key requirements

  • Strong software engineering background, ideally in distributed, real-time systems
  • Experience with containerisation and orchestration (Kubernetes) in production
  • Familiarity with observability tooling (Victoria Metrics, Prometheus, Grafana, OpenTelemetry, SLOs)
  • Strong debugging skills to diagnose and resolve issues under time pressure
  • Proven track record of fault-tolerant, high-availability platforms
  • Experience delivering software in resource-constrained environments with CI/CD and deployment automation
  • Comfort working in a federated model across multiple teams and product streams
  • Focus on continuous improvement and reducing manual intervention
  • Collaborative mindset across cross-functional teams
  • Strong problem-solving and analytical abilities
  • Effective communication under pressure
  • Kubernetes in production
  • Observability tooling: Victoria Metrics, Prometheus, Grafana, OpenTelemetry, SLOs
  • Deployment automation

Posted: September 14th, 2026