Service Design Specialist

Company: ARM
Apply for the Service Design Specialist
Location: Cambridge
Job Description:

Overview

In this role you own the quality and reliability of Arm’s services, applying SRE principles to production and risk-aware design. You’ll champion observable, automated services with strong readiness and reliability standards, building trust across the ecosystem. You will collaborate with engineering and operations to embed SRE practices early, defining meaningful performance targets and ensuring launch-ready services. This is a hands-on, impact-driven opportunity to shape resilient, AI-ready service architectures.

Pay / Benefits

  • accommodations during recruitment
  • hybrid working
  • equal opportunities employer

Responsibilities

  • Define and govern Service Acceptance Criteria aligned to ITIL v4 and SRE
  • Lead risk-based readiness reviews for new services, features, and major releases
  • Govern service transition into production including rollout integrity and rollback capability
  • Embed SLOs, SLIs, SLAs, and XLAs into service design
  • Ensure measurable reliability targets and automation standards aligned with IT operating model
  • Partner with engineering/operations to embed SRE practices early in the lifecycle
  • Validate resilience patterns and address reliability gaps pre-release
  • Embed observability, monitoring, alerting, and health models into service architecture
  • Ensure performance, resilience, automation, and AI-readiness are built into services
  • Maintain service health dashboards and reliability reporting
  • Validate operational docs, runbooks, and support models
  • Proactively identify systemic, architectural, and design risks before incidents

Key requirements

  • ITIL v4 certified with foundation as a minimum
  • 5+ years in IT service design, service transition, or reliability engineering in fast-paced environments
  • Experience embedding SRE principles across end-to-end service lifecycle
  • Experience defining/measuring SLIs/SLOs/SLAs and integrating observability into build
  • Experience in DevOps and CI/CD environments with release/deployment governance
  • Data-driven mindset for using metrics to guide quality decisions
  • collaboration
  • problem solving
  • communication
  • ITIL v4 Foundation
  • SRE principles and reliability practices
  • observability/monitoring/telemetry integration

…

Posted: September 30th, 2026