Data Center IT Operations Manager / Senior IT Operations Engineer -London, UK

Company: Alibaba Cloud
Apply for the Data Center IT Operations Manager / Senior IT Operations Engineer -London, UK
Location: London
Job Description:

  • IT Operations & Availability: Lead the end-to-end Operation&Maintenance of server hardware and network equipment. Develop and execute high-impact O&M strategies to ensure maximum system stability and meet all SLA-defined performance metrics.
  • Incident & Risk Management: Establish robust mechanisms to identify technical risks early. Oversee emergency response plans to minimize downtime and ensure rapid “root cause” resolution for critical server and network issues.
  • Process Standardization: Define and optimize IT Service Management (ITSM) workflows. Ensure all work orders, hardware repairs, and risk closures are documented and executed with precision and compliance.
  • Technical Support and Innovation: Introduce and leverage IT tools to automate routine tasks and improve troubleshooting efficiency. Provide technical guidance to resolve complex onsite hardware or connectivity challenges.
  • Strategic Vendor Management: Lead and manage an outsourced team, taking ownership of the recruitment process, performance evaluation, and training programs to build a high-performing and continuously growing workforce.
  • Stakeholder Coordination: Act as the primary liaison between cross-functional teams and external partners. Coordinate resources effectively to mitigate operational risks and support large-scale business deployments.

Job Requirements

  • Bachelor’s degree in IT, computer science or other relevant domain.
  • Minimum 3 years of leadership experience in IT Operations management, specifically focused on server and network hardware. Prior experience in a large-scale data center environment is highly preferred;
  • Deep knowledge of server architecture, network infrastructure, and IT service delivery principles. Familiar with ITSM (IT Service Management) frameworks.
  • Proven track record of meeting or exceeding rigorous uptime and repair-time SLAs in a fast-paced, high-intensity environment.
  • Demonstrated ability to manage tech teams and coordinate multi-party resources to deliver complex projects on schedule.
  • Exceptional communication and emergency-handling skills. Proactive, highly responsible, and able to maintain composure during critical system outages.
  • Fluency in English and Mandarin Chinese is highly preferred.
  • This is an onsite position requiring physical presence at the facility during business days

#J-18808-Ljbffr…

Posted: August 18th, 2026