Google is hiring a Systems Engineering Manager for Site Reliability Engineering focusing on ML Compute in London. You will lead a team responsible for uptime, availability and performance of core services and drive automation across large-scale infrastructure.
Own end-to-end reliability, mentor engineers, manage on-call rotations across continents, and collaborate with product and platform teams to deliver scalable, fault-tolerant ML compute environments.
#J-18808-Ljbffr…
