Systems Engineering Manager, Site Reliability Engineering, ML Compute

Company: Hackajob Ltd
Apply for the Systems Engineering Manager, Site Reliability Engineering, ML Compute
Location:
Job Description:

Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google’s servicesboth our internally critical and our externally-visible systemshave reliability, uptime appropriate to users’ needs and a fast rate of improvement. Additionally SREs will keep an ever-watchful eye on our systems …

WHJS1_UKTJ

…

Posted: September 24th, 2026