L
Director of Production Engineering
Legion
Contract type
Ongoing
Work mode
100% remote
Experience
Lead / principal · 10+ years
Job description
Key details
- Hire, mentor, and manage a globally-distributed DevOps/SRE engineering team, fostering a culture of ownership, collaboration, and continuous improvement
- Own the reliability and infrastructure roadmap for the AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency
- Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation
- Define and drive engineering OKRs for infrastructure reliability, automation, and security, tracking progress against measurable outcomes
- Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution
- Apply agentic AI workflows to SDLC and DevOps processes, such as automated investigation, remediation, and PR-generation pipelines
- Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil
- Align infrastructure standards, access controls, tooling, and compliance requirements with engineering and IT teams
- Ensure the platform meets the highest standards of security, compliance, and data protection, implementing robust security controls and audit-readiness
- Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals
- Stay current on cloud, DevOps, and security best practices, providing technical guidance and thought leadership
- Company mission
- Information not specified
Primary stack
Core technologies
AWSTerraformPython (Programming Language)Go
Benefits
- Information not specified
Requirements & details
- 8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience
- Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3)
- Demonstrated experience running security operations (SecOps) — vulnerability management, incident response, and remediation of production security issues
- 5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements
- Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation
- Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines
- Hands-on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu)
- Hands-on leadership role with ~20-30% of time contributing directly to architecture, tooling, and incident response
- AWS, EKS, RDS, Datadog, Terraform, CloudFormation, Python, Go, Bash, Git, Linux
- AWS
- Terraform
- Python (Programming Language)
- Go
