All jobs
Save

Senior AI DevOps / LLMOps

TechBiz Global
Contract type
Ongoing
Work mode
100% remote
Experience
Lead / principal · 10+ years

Job description

Key details

  • Automate Build-to-Production processes
  • Design and implement robust CI/CD pipelines tailored for AI, covering model weights, dataset versioning, and application code
  • Develop specialized workflows for PromptOps, ensuring system prompts are version-controlled, tested for regressions, and deployed with the same rigor as traditional code
  • Automate the deployment of Agentic workflows, managing stateful AI interactions and multi-agent handoffs
  • Provision and manage high-performance compute environments (GPU clusters, TPU pods) using Terraform, Pulumi, or Ansible
  • Define and enforce Policy-as-Code for AI endpoints to ensure compliance with security, cost-usage limits, and data residency requirements
  • Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless parity between On-Premises development and Cloud production
  • Architect Progressive Delivery strategies for AI, including Canary releases, Blue-Green deployments, and Shadowing
  • Build Evaluation-in-the-Loop gates within the pipeline to automatically test for bias, hallucination, and performance degradation before a release
  • Implement A/B testing frameworks specifically designed for LLM outputs and agentic behavior
  • Establish deep observability into Inference Endpoints, tracking metrics like tokens-per-second, latency, and drift in model accuracy
  • Integrate feedback loops that capture production edge cases to feed back into the training and fine-tuning pipelines
  • Company mission
  • Information not specified

Primary stack

Core technologies

KubernetesTerraformAWSAzure

Benefits

  • Full time and remote job

Requirements & details

  • 10+ years in DevOps, SRE, or Cloud Engineering
  • 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs from notebook to production
  • Proven experience managing Hybrid Cloud environments (e.g., AWS/Azure + Private Data Center)
  • Advanced Kubernetes (K8s) skills, specifically with KubeFlow, Ray, or NVIDIA Triton
  • Expertise in GitHub Actions/GitLab CI, and Terraform or Pulumi
  • Experience with Weights & Biases, MLflow, LangSmith, or Arize Phoenix
  • Understanding of GPU virtualization, CUDA drivers, and on-premises hardware management
  • Familiarity with Open Policy Agent (OPA) and secret management (Vault)
  • Fluent English is needed
  • Kubernetes, KubeFlow, Ray, NVIDIA Triton, GitHub Actions, GitLab CI, Terraform, Pulumi, Ansible, Weights & Biases, MLflow, LangSmith, Arize Phoenix, Open Policy Agent, Vault, AWS, Azure, CUDA, LLMs, AI Agents
  • Kubernetes
  • Terraform
  • AWS
  • Azure

Apply