Site Reliability Engineer - Department for Business and Trade - SEO
Government Digital & Data -
Location
Belfast, Birmingham, Cardiff, Darlington, Edinburgh, London, Salford
About the job
Job summary
The Department for Business and Trade (DBT) has a clear mission - to grow the economy. Our role is to help businesses invest, grow and export to create jobs and opportunities right across the country. We do this in three ways.
Firstly, we help to build a strong, competitive business environment, where consumers are protected and companies rewarded for treating their employees properly.
Secondly, we open international markets and ensure resilient supply chains. This can be through Free Trade Agreements, trade facilitation and multilateral agreements.
Finally, we work in partnership with businesses every day, providing advance, finance and deal-making support to those looking to start up, invest, export and grow.
The Digital, Data and Technology (DDaT) directorate develops and operates tools and services to support us in this mission.
Job description
As a Site Reliability Engineer, you will pro-actively engage development teams and use initiative to develop the tools for their job, including application performance monitoring, exception, log and metrics aggregation, dashboards, and declarative CI/CD (continuous integration/continuous delivery) pipelines.
You’ll work alongside senior members of the team to pro-actively engage product teams about service-level indicators, objectives, and error budgets. You’ll contribute towards building and scaling our shared tooling for the platform and participate in an on-call rota.
Our tech stack includes:
- AWS and Azure
- GitHub Actions and AWS CodePipelines/CodeBuild
- Terraform
- Containerisation such as Docker, Elastic Container Service (ECS), Elastic Container Registry (ECR) and Kubernetes
- ElasticSearch/OpenSearch
- Python and Django framework
- PostgreSQL as a service (Amazon RDS)
- Datadog, Logstash
- Redis/Elasticache
Person specification
You should be able to demonstrate:
- Cloud experience with either Amazon Web Services, Azure or Google Cloud.
- Exposure to building code-defined and reliable infrastructure on top of cloud computing systems (e.g. Terraform, CloudFormation, Pulumi).
- Working knowledge in one or more programming languages, writing clean and effective code.
- Working knowledge of Linux/Unix fundamentals and TCP/IP networking.
- Ability to see user impact in the infrastructure or application changes, including a drive to improve the end user experience at every turn.
- Strong communication skills when dealing with both technical and non-technical stakeholders.