Distinguished Engineer (Head of Service Reliability Engineering) - HMRC - SCS1
Government Digital & Data -
Location
Bristol, Leeds, London, Newcastle-upon-Tyne, Telford
About the job
Job summary
This is a rare opportunity to shape reliability engineering across one of the UK’s largest and most important digital estates.
Enterprise Live Services (ELS) at HMRC is evolving into a modern, engineering-led organisation - enabling HMRC to deliver change faster while strengthening reliability, resilience and trust.
As Head of Service Reliability Engineering, you will lead a distinct, independent capability spanning services, products and platforms. You will make reliability, operability and observability integral to engineering from the outset.
With enterprise-wide reach and influence, you will set the direction for SRE, raise service maturity and build a lasting capability that enables teams to innovate safely and operate with confidence.
Job description
The Head of Service Reliability Engineering is accountable for defining, embedding, and continuously raising service reliability engineering capability across HMRC.
Sitting within Chief Engineering & Platform Office (CEPO) and Enterprise Live Services (ELS), the role leads the independent SRE tower, operating as a centre of excellence for reliability, operability, resilience, and observability. The role does not own delivery, platforms, or live operations. Instead, it coaches, challenges, and enables product, platform, and ELS teams to meet enterprise reliability and service maturity expectations using objective, evidence‑based practices.
The role provides strategic leadership and technical authority to ensure that:
- Reliability and operability are designed into services from inception
- Assurance is grounded in observable service behaviour and maturity evidence
- Enterprise standards enable autonomy and speed, rather than acting as delivery gates
- HMRC can demonstrate confidence, auditability, and resilience at scale
Person specification
Technical Leadership: Provides senior technical leadership in enterprise-scale Reliability Engineering, with the credibility to influence architecture, set standards and deliver reliability outcomes in complex, regulated environments.
Site Reliability Engineering (SRE): Leads the adoption of SRE practices by using SLOs, error budgets and reliability metrics to drive measurable improvements in service performance and operational decision-making.
Observability: Establishes and governs enterprise observability capabilities, using telemetry, dependency mapping and modern monitoring platforms to improve operational insight and decision quality.
Resilience and Recovery: Designs and embeds resilience practices that validate recovery capabilities, identify operational risks and strengthen service continuity through testing and evidence-based improvements.
Secure-by-Design Operations: Integrates security and reliability disciplines by using security intelligence, risk indicators and close collaboration with cyber teams to enhance service operability and resilience.
Cloud and Platform Engineering: Applies deep expertise in cloud-native and platform-based technologies to assess and mitigate reliability risks associated with architecture, platforms and continuous delivery environments.
Service Maturity and Assurance: Develops and applies service maturity frameworks that use evidence-based assessment to inform risk, investment and operational decisions rather than relying on compliance measures.
Executive Stakeholder Management: Communicates complex technical risks clearly to senior stakeholders and uses objective evidence to influence strategic decisions at Director and Executive level.
What an Outstanding Candidate Would Also Bring (Desirable): Brings experience operating in highly regulated environments, shaping engineering standards across organisations, building reliability communities of practice and supporting modern service management approaches.