W
Infrastructure engineer (UK)
Writer
Contract type
Ongoing
Work mode
100% remote
Experience
Senior · 5+ years
Job description
Key details
- Bring deep focus to one substantial infrastructure initiative at a time, moving between SRE, DevOps, Infrastructure, and Platform work as leverage shifts
- Challenge the status quo and remove toil before adding features by automating operational tasks and infrastructure management with Python or Go
- Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and supporting cloud and AI tooling
- Run agents in your daily loop (Claude Code, Droid, Codex, internal skills) to investigate incidents, draft Terraform/Helm changes, write runbooks, scaffold tooling, and review PRs
- Encode recurring infrastructure tasks as internal skills any teammate (human or agent) can pick up and run
- Lead incident response, post-mortems, and root-cause analyses, tracing failures to the underlying problem and applying learnings back into the architecture
- Own the reliability, performance, and efficiency of core services end-to-end, defining and upholding SLOs and error budgets, and carrying the on-call pager
- Balance this week's critical work with the 6–12-month platform direction, shaping multi-year observability, cost, and reliability investments
- Operate cross-functionally with product, security, and engineering peers, providing expert guidance on system design for reliability, performance, and scalability from conception through launch
- Connect the infrastructure agenda to product and revenue context, and disagree with evidence, not volume
- Company mission
- Information not specified
Primary stack
Core technologies
Python (Programming Language)GoAWSGoogle Cloud Platform (GCP)AzureKubernetesTerraform
Benefits
- Information not specified
Requirements & details
- 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company
- Experience running containerisation in production (a real cluster, not a lab), with experience in Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)
- Good proficiency in Python or Go for automation and tooling
- AI is part of how you ship, not a thing you've read about — agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop, you've built or adopted AI-assisted workflows others now use, and you have strong opinions on where it is unreliable
- Python, Go, AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Claude Code, Droid, Codex
- Python (Programming Language)
- Go
- AWS
- Google Cloud Platform (GCP)
- Azure
- Kubernetes
- Terraform
