Staff Slurm Cluster & HPC Engineer @Bitdeer Technologies Group
Engineering
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1mth ago

[Hiring] Staff Slurm Cluster & HPC Engineer @Bitdeer Technologies Group

1mth ago - Bitdeer Technologies Group is hiring a remote Staff Slurm Cluster & HPC Engineer. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability on both bare-metal and VM-based GPU nodes, and will lead our adoption of the Slinky operator stack (slurm-operator, and slurm-bridge where it fits) so that Slurm and Kubernetes workloads can share the same GPU pool. The role is deeply hands-on, customer-facing during onboarding and escalations, and sets the engineering standard the rest of the platform team builds on.

Key Responsibilities

  • Slurm cluster architecture and lifecycle
    • Design, deploy, and operate production Slurm clusters on bare metal and VMs.
    • Ensure slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
  • Topology-aware scheduling for GPU fabrics
    • Model the physical fabric in topology.conf.
    • Prove placement quality with NCCL bandwidth and multi-node training validation.
  • Multi-tenant scheduling policy
    • Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits.
    • Enforce fail-closed defaults.
  • Slinky on Kubernetes
    • Lead implementation of the Slinky slurm-operator.
    • Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods.
  • Elastic capacity between Slurm and Kubernetes
    • Use Slurm cloud and power-save mechanisms together with fleet automation.
  • Container and job runtime
    • Operate Pyxis/Enroot and OCI/containerd job paths.
    • Support MPI/PMIx, module/Spack environments, and customer-supplied images.
  • Cluster health and reliability engineering
    • Build the passive and active health-check system.
    • Own burn-in and acceptance testing for every new rack.
  • Automation and infrastructure as code
    • Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning.
  • Observability, accounting, and billing integration
    • Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy.
    • Configure AccountingStorageTRES and TRESBillingWeights.
  • Technical leadership and customer engagement
    • Write runbooks and tenant-facing documentation.
    • Onboard and support enterprise customers.
    • Mentor platform engineers on Slurm and HPC scheduling practice.

Qualifications

  • 8+ years in HPC, systems, or cloud infrastructure engineering.
  • 4+ years operating production Slurm clusters at 100+ GPU-node scale.
  • Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, etc.
  • Strong GPU and fabric fundamentals.
  • Production Kubernetes experience.
  • Experience delivering both bare-metal and virtualized compute.
  • Working knowledge of parallel and shared storage for AI workloads.
  • Proficient in Python and Bash for cluster automation.
  • Multi-tenant security discipline.
  • Clear written and verbal communication in English.

Benefits

  • Equal employment opportunities in accordance with country, state, and local laws.
  • No discrimination against employees or applicants based on various conditions.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Staff Slurm Cluster & HPC Engineer @Bitdeer Technologies Group
Engineering
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,809+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later