All jobs
Save

Director of Production Engineering

Legion
Contract type
Ongoing
Work mode
100% remote
Experience
Lead / principal · 10+ years

Job description

Key details

  • Hire, mentor, and manage a globally-distributed DevOps/SRE engineering team, fostering a culture of ownership, collaboration, and continuous improvement
  • Own the reliability and infrastructure roadmap for the AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency
  • Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation
  • Define and drive engineering OKRs for infrastructure reliability, automation, and security, tracking progress against measurable outcomes
  • Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution
  • Apply agentic AI workflows to SDLC and DevOps processes, such as automated investigation, remediation, and PR-generation pipelines
  • Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil
  • Align infrastructure standards, access controls, tooling, and compliance requirements with engineering and IT teams
  • Ensure the platform meets the highest standards of security, compliance, and data protection, implementing robust security controls and audit-readiness
  • Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals
  • Stay current on cloud, DevOps, and security best practices, providing technical guidance and thought leadership
  • Company mission
  • Information not specified

Primary stack

Core technologies

AWSTerraformPython (Programming Language)Go

Benefits

  • Information not specified

Requirements & details

  • 8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience
  • Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3)
  • Demonstrated experience running security operations (SecOps) — vulnerability management, incident response, and remediation of production security issues
  • 5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements
  • Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation
  • Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines
  • Hands-on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu)
  • Hands-on leadership role with ~20-30% of time contributing directly to architecture, tooling, and incident response
  • AWS, EKS, RDS, Datadog, Terraform, CloudFormation, Python, Go, Bash, Git, Linux
  • AWS
  • Terraform
  • Python (Programming Language)
  • Go

Apply