G
Software Golang Engineer (Slurm)
Gcore
Contract type
Ongoing
Work mode
Hybrid
Experience
3 years
Job description
Key details
- Submit and debug Slurm workloads using sbatch, srun, squeue, and sinfo in production environments
- Build production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops in Go
- Preserve traditional Slurm cluster behavior while running underlying infrastructure on Kubernetes
- Diagnose performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
- Treat Slurm as a customer-facing platform with a product mindset and strong customer empathy
- Take end-to-end ownership of complex distributed-system challenges with excellent communication
- Company mission
- Information not specified
Primary stack
Core technologies
GoKubernetes
Benefits
- Competitive compensation
- Flexible working hours and hybrid or remote options, depending on your role
- Work from anywhere in the world for up to 45 days per year
- Private medical insurance for you and your family (may vary by location)
- Extra paid vacation and sick leave days (may vary by location)
- Support for life's important moments and celebrations
- Language courses to help you connect and grow
- Modern, welcoming offices with snacks, drinks, and entertainment (may vary by location)
- Team sports and social activities (may vary by location)
Requirements & details
- Hands-on experience using Slurm in production from a user's perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo
- Strong proficiency in Go, with experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops
- Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes
- Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
- A product mindset and strong customer empathy, treating Slurm as a customer-facing platform rather than simply another system daemon
- Excellent communication skills and the ability to take end-to-end ownership of complex distributed-system challenges
- Experience operating large-scale HPC or GPU clusters for external customers
- Experience with PyTorch distributed training and other large-scale AI/ML frameworks
- Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure
- Experience building unified job-submission workflows across Kubernetes and Slurm
- Experience in GPU-cloud or HPC product engineering environments
- Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects
- Go, Kubernetes, Slurm, PyTorch, InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph
- Go
- Kubernetes
