Lead Software Engineer – Software Reliability

Company: JP Morgan Chase
Apply for the Lead Software Engineer – Software Reliability
Location: Glasgow
Job Description:

Overview

In this role you design, build, and operate scalable, resilient systems using Python within JPMorgan Chase’s product team. You apply SRE principles to boost availability, performance, and reliability while helping establish engineering standards and security. You collaborate with engineering and product partners to troubleshoot and optimize production services, fostering an inclusive, high-performing culture. Your work drives measurable improvements in system health and reliability at scale, using modern practices and data-driven insights. This is a mission-focused opportunity to shape dependable software across critical financial platforms.

Responsibilities

  • Design and develop scalable, resilient systems using Python with SRE concepts
  • Build secure, high-quality production code and maintain critical algorithms
  • Contribute to architecture and design artifacts ensuring constraints are met
  • Analyze data and create visualizations/reports to improve software and systems
  • Identify data patterns to improve coding hygiene and system architecture
  • Implement reliability practices: monitoring, alerting, automated recovery
  • Define and measure SLIs/SLOs to monitor system health
  • Conduct chaos engineering to test resiliency and identify weaknesses
  • Perform performance testing with tools like JMeter for scalability
  • Collaborate with product teams to enhance reliability and participate in communities of practice
  • Foster an inclusive team culture and drive post-incident reviews and root cause analysis

Key requirements

  • Hands-on system design and Python development experience
  • Experience developing, debugging, and maintaining code in a large corporate environment
  • Knowledge of SDLC, AWS, troubleshooting, resiliency, and automation
  • Understanding of agile methodologies, CI/CD, application resiliency, and security
  • Experience with AI/ML contexts and MongoDB
  • Familiarity with reliability engineering concepts: monitoring, alerting, automated recovery
  • Ability to implement system health checks and performance metrics
  • Understanding of SRE principles, SLIs/SLOs, and incident response
  • Experience with chaos engineering and performance testing (e.g., JMeter)
  • collaborative
  • inclusive
  • effective communication
  • Python
  • AWS
  • CI/CD

…

Posted: October 2nd, 2026