Senior Infrastructure Operations Engineer (DevOps/Platform) - HMRC - SEO
Government Digital & Data -
Location
Leeds, Manchester, Telford
About the job
Job summary
Discover a career in your hands at HMRC. Whether you're seeking purpose, growth, or a workplace that gives you a true sense of belonging, hear from some of our employees as they share their story about what it’s really like to work at HMRC.
Visit our YouTube channel to watch the full series and come and discover your potential.
Are you passionate about cloud platform operations, security, and automation?
Do you have experience in addressing risks, incidents and service quality across technical teams?
Do you want to play a key role in delivering reliable, secure, and efficient platforms that provide services for millions of citizens?
We are looking for a resourceful individual who can ensure our cloud platforms are robust, secure, and optimal. From monitoring performance and managing change to enabling automation and driving standardisation, you’ll help keep HMRC’s mission-critical cloud platforms running at their best.
Within HMRC’s Chief Digital & Information Group (CDIO), specifically in the Enterprise Cloud Services (ECS) team we are redefining our offerings and growing the team of outstanding people to improve the Cloud Centre of Excellence. We are already a diverse team of 90+ technologists, creating a dynamic and inclusive working environment whose skills cover architecture, platform development, service design, platform operations and governance.
Important - Travel to Telford is required as part of this role, and 60% of your working time will need to be office based.
Additional Security Information:
This role requires the successful candidates to hold or be willing to hold Security Check (SC) clearance.
Job description
As a Senior Infrastructure Operations Engineer (Platform Operations Engineer) within HMRC’s Enterprise Cloud Services (ECS), you will ensure the reliability, security, and efficiency of platforms hosting HMRC’s mission-critical services. You will undertake day-to-day operations and drive continual improvement across our cloud platforms, maintaining a high level of availability through proactive monitoring and rapid incident response.
Security and compliance will be central to your role, applying access controls, patching, and vulnerability management to meet industry standards. You will apply change and configuration processes to minimise risk and maintain accurate system data, while using metrics, logging, and automation to optimise performance, reduce manual effort, and control costs.
Working closely with service desks and platform engineering teams, you will resolve incidents, support users, and maintain CI/CD pipelines and Infrastructure as Code. By enabling observability and feedback loops, you will help ensure our platforms continue to run reliably. You will also document procedures, standardise practices, and embrace good operational practises to strengthen resilience and consistency across the organisation.
Person specification
Your responsibilities will include:
Platform Operations & Reliability
- Own the operational health of our Cloud platforms.
- Monitor, troubleshoot, and resolve incidents across our Cloud platforms.
- Drive root-cause analysis and implement permanent fixes.
Automation & Infrastructure as Code
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Remove manual intervention through an automation-first mindset.
- CI/CD with GitLab.
- Design, implement, and maintain CI/CD pipelines in GitLab to automate deployments and configuration changes.
- Iterate on automated testing, security scanning, and compliance checks within pipelines.
- Collaborate across our teams on reusable GitLab CI templates and standardising build/release processes.
- Optimise pipeline performance and troubleshoot issues.
Security & Compliance
- Apply cloud security according to industry good practice.
- Support audits and maintain alignment with compliance frameworks.
- Manage patching and vulnerabilities through automation.
Performance & Cost Optimisation
- Analyse value/cost/performance to optimise spend and capacity.
- Provide guidance to development teams on efficient resource usage.
Collaboration & Leadership
- Mentor engineers and share best practices across Platform teams.
- Partner with platform development teams to ensure smooth releases.
Mentoring & Technical Leadership
- The ideal candidate will demonstrate a proactive approach to mentoring and supporting junior team members, fostering a culture of continuous learning and growth.
- They will actively contribute to the technical leadership of the team by sharing knowledge, offering constructive feedback, and guiding others through complex problem-solving.
- Their ability to lead by example, communicate clearly, and encourage collaboration will be key to building a resilient and high-performing engineering culture.
Critical Thinking & Technical Mindset
- We’re looking for someone who brings a strong critical thinking mindset to their technical work - someone who doesn’t just follow patterns, but questions assumptions, evaluates options, and makes informed decisions.
- The successful candidate will approach challenges with curiosity and analytical rigour, balancing innovation with pragmatism.
- They will be comfortable navigating ambiguity, identifying root causes, and proposing solutions that are both technically sound and aligned with strategic goals.
Essential Criteria:
- Experience operating in AWS, or Azure.
- Proven incident & problem management, and root-cause analysis capability.
- Proficiency with Linux systems administration.
- Capability in Infrastructure as Code (Terraform preferred).
- Knowledge of GitLab CI/CD, including pipeline design, runners, and secrets management.
- Scripting skills (e.g. Python, Go, or Bash).
List Desirable Criteria:
- Experience with Monitoring & Observability tooling.
- Exposure to configuration management tools.
- Awareness of FinOps and cloud cost management.
- Cloud related certifications.
- Familiarity with the principles behind ITIL service operations.