Software Engineer III – AI/ML Platform Reliability

Company: JP Morgan Chase
Apply for the Software Engineer III – AI/ML Platform Reliability
Location: Glasgow
Job Description:

Overview

In this role you will strengthen the reliability and scalability of AI/ML platforms within a global financial services tech organization. You will collaborate across teams to deliver trusted, production-grade systems and tooling that support large-scale AI capabilities. You’ll own reliability requirements, contribute to observability and security, and tackle complex production problems in a fast-growing environment. This opportunity lets you shape how the firm operationalizes AI with a focus on resilience and secure delivery.

Responsibilities

  • Enhance reliability and scalability of AI/ML platforms and applications to meet growing demand
  • Own non-functional requirements and develop tooling for observability, security, resilience, and operations excellence
  • Build and maintain scalable infrastructure for deployment and operation of large-scale AI platforms and apps
  • Foster cross-functional relationships and deliver solutions to user problems
  • Participate in on-call rotations, troubleshoot production issues, and take ownership of problems
  • Develop and review secure, high-quality production code and assist peers with debugging
  • Automate remediation of recurring issues to improve system stability
  • Leverage enterprise AI-assisted development tools to improve code quality, delivery speed, and productivity while ensuring secure coding and peer review

Key requirements

  • Formal training or certification in software engineering concepts
  • Hands-on experience delivering system design, application development, testing, and operational stability
  • Advanced proficiency in Python
  • Experience across the Software Development Life Cycle
  • Experience with infrastructure-as-code and cloud-native delivery (Terraform, containers, Kubernetes, CI/CD, automated deployment)
  • Experience designing and developing large-scale distributed systems and cloud-native architectures
  • Experience building large-scale infrastructure in Google Cloud, AWS, or Azure with Terraform
  • Extensive experience implementing observability (Open Telemetry, Dynatrace, Grafana, etc)
  • Strong problem-solving in complex systems
  • Experience using AI-assisted development tools and validating AI outputs
  • Understanding of responsible AI use in engineering workflows
  • Strong ownership and proactive, self-motivated work style
  • Cross-functional collaboration
  • Strong problem-solving and troubleshooting
  • Ownership and urgency
  • Python
  • Software Development Life Cycle
  • Terraform

…

Posted: October 1st, 2026