All jobs
Save

Software Golang Engineer (Slurm)

Gcore
Contract type
Ongoing
Work mode
Hybrid
Experience
3 years

Job description

Key details

  • Submit and debug Slurm workloads using sbatch, srun, squeue, and sinfo in production environments
  • Build production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops in Go
  • Preserve traditional Slurm cluster behavior while running underlying infrastructure on Kubernetes
  • Diagnose performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
  • Treat Slurm as a customer-facing platform with a product mindset and strong customer empathy
  • Take end-to-end ownership of complex distributed-system challenges with excellent communication
  • Company mission
  • Information not specified

Primary stack

Core technologies

GoKubernetes

Benefits

  • Competitive compensation
  • Flexible working hours and hybrid or remote options, depending on your role
  • Work from anywhere in the world for up to 45 days per year
  • Private medical insurance for you and your family (may vary by location)
  • Extra paid vacation and sick leave days (may vary by location)
  • Support for life's important moments and celebrations
  • Language courses to help you connect and grow
  • Modern, welcoming offices with snacks, drinks, and entertainment (may vary by location)
  • Team sports and social activities (may vary by location)

Requirements & details

  • Hands-on experience using Slurm in production from a user's perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo
  • Strong proficiency in Go, with experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops
  • Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes
  • Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
  • A product mindset and strong customer empathy, treating Slurm as a customer-facing platform rather than simply another system daemon
  • Excellent communication skills and the ability to take end-to-end ownership of complex distributed-system challenges
  • Experience operating large-scale HPC or GPU clusters for external customers
  • Experience with PyTorch distributed training and other large-scale AI/ML frameworks
  • Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure
  • Experience building unified job-submission workflows across Kubernetes and Slurm
  • Experience in GPU-cloud or HPC product engineering environments
  • Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects
  • Go, Kubernetes, Slurm, PyTorch, InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph
  • Go
  • Kubernetes

Apply