Lead Engineer (Production)

Company: Tria Recruitment
Apply for the Lead Engineer (Production)
Location: Glasgow
Job Description:

Job Description n Location: London, onsite 3 days per week (Sheffield as an alternative) n Rate: £tbd/day inside IR35 n Duration: 6 months+ n Are you a Senior Observability Engineer/SRE Lead, with demonstrable experience of assessing and defining observability and monitoring roadmaps within enterprise scale environments? If so, apply now for this new contract opportunity. n The Lead Observability Engineer/SRE Lead will be required to assess a complex hybrid estate, understand how services, platforms, infrastructure and networks should be monitored, and work across multiple internal teams, partners and suppliers to build a consolidated view of existing telemetry, monitoring and alerting capabilities. n As well as short term tactical objectives, the role will also be focussed on longer term strategic ones. n Responsibilities of the Lead Observability Engineer/SRE Lead will be to: n n Discover and document existing telemetry sources, monitoring tools, dashboards and ownership n Work with technical teams and suppliers to gain access to telemetry n Deliver meaningful dashboards and service health views n Identify gaps in telemetry, monitoring and alerting coverage and implement pragmatic improvements n Define health indicators for critical business journeys n Introduce modern observability practices where practical, including SLIs, SLOs and a roadmap towards burn-rate alerting n Develop a roadmap for OpenTelemetry adoption n Assess options for a centralised telemetry platform, including Grafana Cloud n Evaluate tooling rationalisation opportunities, operating costs and telemetry economics n Define an observability target architecture, standards and implementation roadmap n n The successful Lead Observability Engineer/SRE Lead will demonstrate the following: n n Proven experience leading enterprise-scale observability initiatives n Strong hands-on expertise with Grafana, Grafana Cloud, OpenTelemetry and modern telemetry pipelines n Deep understanding of metrics, logs, traces, distributed tracing, alerting, SLIs, SLOs and error-budget concepts n Experience designing observability solutions across cloud PaaS, IaaS, on-premises, Legacy and third-party hosted platforms n Strong knowledge of Azure observability tooling, including Azure Monitor, Log Analytics and Application Insights n n If this sounds like you, please apply to find out more. n Lead Observability Engineer/SRE Lead/Lead Site Reliability Engineer

…

Posted: October 6th, 2026