All jobs
Save

Infrastructure engineer (UK)

Writer
Contract type
Ongoing
Work mode
100% remote
Experience
Senior · 5+ years

Job description

Key details

  • Bring deep focus to one substantial infrastructure initiative at a time, moving between SRE, DevOps, Infrastructure, and Platform work as leverage shifts
  • Challenge the status quo and remove toil before adding features by automating operational tasks and infrastructure management with Python or Go
  • Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and supporting cloud and AI tooling
  • Run agents in your daily loop (Claude Code, Droid, Codex, internal skills) to investigate incidents, draft Terraform/Helm changes, write runbooks, scaffold tooling, and review PRs
  • Encode recurring infrastructure tasks as internal skills any teammate (human or agent) can pick up and run
  • Lead incident response, post-mortems, and root-cause analyses, tracing failures to the underlying problem and applying learnings back into the architecture
  • Own the reliability, performance, and efficiency of core services end-to-end, defining and upholding SLOs and error budgets, and carrying the on-call pager
  • Balance this week's critical work with the 6–12-month platform direction, shaping multi-year observability, cost, and reliability investments
  • Operate cross-functionally with product, security, and engineering peers, providing expert guidance on system design for reliability, performance, and scalability from conception through launch
  • Connect the infrastructure agenda to product and revenue context, and disagree with evidence, not volume
  • Company mission
  • Information not specified

Primary stack

Core technologies

Python (Programming Language)GoAWSGoogle Cloud Platform (GCP)AzureKubernetesTerraform

Benefits

  • Information not specified

Requirements & details

  • 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company
  • Experience running containerisation in production (a real cluster, not a lab), with experience in Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)
  • Good proficiency in Python or Go for automation and tooling
  • AI is part of how you ship, not a thing you've read about — agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop, you've built or adopted AI-assisted workflows others now use, and you have strong opinions on where it is unreliable
  • Python, Go, AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Claude Code, Droid, Codex
  • Python (Programming Language)
  • Go
  • AWS
  • Google Cloud Platform (GCP)
  • Azure
  • Kubernetes
  • Terraform

Apply