Senior Lead Software Engineer – LLM Ops Platform Reliability

Company: JP Morgan Chase
Apply for the Senior Lead Software Engineer – LLM Ops Platform Reliability
Location: Glasgow
Job Description:

Overview

In this role you will design and operate production-grade AI infrastructure at enterprise scale, focusing on reliable LLM serving, cloud and Kubernetes deployments, and deep observability. You’ll own end-to-end reliability, performance, and cost-efficiency of AI inference platforms, partnering across teams to improve stability and secure software delivery. You’ll solve hard production problems and drive continuous improvement in a fast-paced, risk-aware environment. This role offers the opportunity to shape AI infrastructure that modernizes traditional SRE practices.

Responsibilities

  • Design, develop, and deliver secure, production-quality software for AI infrastructure
  • Build backend services and APIs enabling reliable AI infrastructure in production
  • Operate and scale LLM serving infrastructure including hosting, routing, batching, and caching
  • Deploy and lifecycle-manage LLMs on cloud and on-prem GPUs using IaC and CD pipelines
  • Implement observability across logs, metrics, and traces with dashboards and alerting
  • Tune GPU capacity, autoscaling, and cost-efficiency using optimization techniques
  • Lead reliability engineering for LLM endpoints including capacity planning and incident response
  • Participate in on-call rotations, incident triage, and post-incident analyses
  • Identify recurring issues and automate remediation to improve stability and developer experience
  • Build and maintain multi-agent systems with orchestration and workflow control
  • Drive adoption of AI-assisted engineering practices with standardized validation and code quality processes

Key requirements

  • Hands-on system design, development, testing, and operations in production
  • Advanced Python proficiency for production services
  • Automation and continuous delivery experience
  • Cloud infrastructure and IaC tooling experience
  • Strong site reliability engineering knowledge (incident management, runbooks, reliability patterns)
  • Observability experience across metrics, logs, and traces
  • Kubernetes and container orchestration experience
  • Experience hosting and serving LLMs on cloud and local GPU environments
  • Knowledge of LLM reliability and risk considerations (latency, throughput, versioning, logging)
  • Experience with enterprise AI-assisted software development tools and evaluating AI outputs for correctness and security
  • Understanding of responsible AI use in engineering workflows and security expectations
  • Collaborative mindset
  • Strong incident management and communication
  • Ownership and drive for reliability improvements
  • Python production services
  • Automation and CD/CI pipelines
  • Cloud platforms and IaC tooling

Posted: September 14th, 2026