Software Golang Engineer (Slurm)

Status
Open
Remote policy
Remote
Employment type
Not stated
Salary
Not stated
Categories
Software-Engineer, Systems-Engineering, HPC-Engineering, Platform-Engineering, Site-Reliability-Engineering, Senior-Golang-Software-Engineer, HPC-Software-Engineer, High-Performance-Computing-(HPC)-Software-Engineer, GoLang-Engineer, Senior-Golang-Developer
Tech
grafanakubernetesgo
Source
himalayas
First observed
2026-08-24 03:28 UTC
Last seen
2026-08-24 03:28 UTC
Source claims posted
2026-08-24 03:01 UTC
Consecutive misses
1 of 10

What the posting says

What You’ll Do

Design and build a managed Slurm service on Kubernetes

Write clean, reliable, and maintainable Go code

Develop scheduling and orchestration capabilities for GPU-intensive and distributed workloads

Build observability and automated remediation for GPU, node, network, and control-plane failures using VictoriaMetrics, Grafana, DCGM

What We're Looking For

Hands-on experience using Slurm in production from a user’s perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo

Strong proficiency in Go, with experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops

Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes

Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems

A product mindset and strong customer empathy, treating Slurm as a customer-facing platform rather than simply another system daemon

Excellent communication skills and the ability to take end-to-end ownership of complex distributed-system challenges

Nice to Have

Experience operating large-scale HPC or GPU clusters for external customers

Experience with PyTorch distributed training and other large-scale AI/ML frameworks

Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure

Experience building unified job-submission workflows across Kubernetes and Slurm

Experience in GPU-cloud or HPC product engineering environments

Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects

Benefits

At Gcore, we want you to do your best work and enjoy the journey. Our benefits are designed to support your growth, well-being, and life beyond work:

Competitive compensation

Flexible working hours and hybrid or remote options, depending on your role

Work from anywhere in the world for up to 45 days per year

Private medical insurance for you and your family*

Extra paid vacation and sick leave days*

Support for life’s important moments and celebrations

Language courses to help you connect and grow

Modern, welcoming offices with snacks, drinks, and entertainment*

Team sports and social activities*

*Benefits may vary depending on your location.

Equal Opportunity Employer

We provide equal opportunity to all applicants without regard to race, color, religion, sex, sexual orientation, age, gender identity, gender expression, national origin, disability, or any other legally protected characteristics.

The world’s digital experiences run on something invisible: the infrastructure and software that keep them fast, reliable, and secure. At Gcore, you’ll help design and deliver that foundation for an AI-driven world.

We’re a global provider of infrastructure and software solutions for AI, cloud, network, and security, powering everything from real-time communication and streaming to enterprise AI and secure web applications. With 210+ edge locations, 50+ cloud regions, and thousands of GPUs, your work here can reach users and businesses across the globe.

You’ll collaborate with leading technology partners such as Intel, NVIDIA, Dell, and Equinix, and work on platforms that power digital products used around the world. Our vision is simple: to connect the world to AI, anywhere, anytime.

Want to work on technology that goes beyond a single product or industry? Join a global team of 550+ professionals building infrastructure and software that supports the entire digital ecosystem.

Originally posted on Himalayas

Quality

Completeness: 65%

Not enough history yet to judge honesty signals.

Timeline

  1. *
    #314074 2026-08-24 03:28 UTC
    Published
  2. o
    #316336 2026-08-24 05:29 UTC
    Not seen
    Miss 1 in a row