Position:
Staff Slurm Cluster & HPC Engineer
Company:
Bitdeer
Location:
Singapore, SG
Employment type:
Full-Time
Short Summary:
Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud, committed to providing comprehensive Bitcoin mining solutions and advanced cloud capabilities.
Responsibilities:
- Design, deploy, and operate production Slurm clusters on bare metal and VMs.
- Model physical fabric for topology-aware scheduling for GPU fabrics.
- Own multi-tenant scheduling policy and enforce fail-closed defaults.
- Lead implementation of Slinky slurm-operator and evaluate slurm-bridge.
- Use Slurm cloud and power-save mechanisms for elastic capacity.
- Operate container and job runtime with correct configurations.
- Build health-check systems for cluster reliability.
- Deliver clusters through Terraform/Ansible and ensure reproducibility.
- Instrument observability and integrate accounting and billing.
- Write documentation, onboard customers, and mentor engineers.
Requirement:
- 8+ years in HPC, systems, or cloud infrastructure engineering.
- 4+ years operating production Slurm clusters at 100+ GPU-node scale.
- Deep hands-on Slurm expertise and strong GPU and fabric fundamentals.
- Production Kubernetes experience and knowledge of Slurm-on-Kubernetes stacks.
- Experience with bare-metal and virtualized compute environments.
- Working knowledge of parallel and shared storage for AI workloads.
- Proficient in Python and Bash; Go experience is a plus.
- Multi-tenant security discipline and clear communication skills in English.
Benefits:
- A culture that values authenticity and diversity.
- An inclusive environment with open workspaces.
- Fast-growing company with networking opportunities.
- Ability to contribute directly to the digital asset industry.
- Personal accountability, autonomy, and learning opportunities.
- Attractive welfare benefits and developmental opportunities.