Position:
Data Centre Infrastructure Engineer
Company:
Bitdeer
Location:
Malaysia, Cyberjaya / Johor Bahru
Employment type:
Full-Time
Short Summary:
Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud, committed to providing comprehensive solutions for its customers.
Responsibilities:
- Responsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.
- Perform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure.
- Monitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.
- Conduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.
- Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
- Troubleshoot hardware and infrastructure issues, including server failures, GPU errors, storage issues, network connectivity problems, switch failures, and cabling faults.
- Execute incident response procedures and provide timely escalation and resolution according to operational standards.
- Prepare shift handover reports and maintain operation documents, SOPs, and incident reports.
- Work closely with engineering, network, and infrastructure teams to support new deployments and ongoing operation improvements.
- Participate in a two-shift rotation schedule, including night shifts, weekends, and holidays as required.
Requirement:
- Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
- Basic understanding of Data Center infrastructure and server hardware architecture.
- Familiarity with systems such as NVIDIA GB200 Cluster, NVIDIA GB300 Cluster, GPU Servers, x86 Servers, Storage Servers, Ethernet and InfiniBand Networks.
- Knowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
- Basic Linux administration skills, including system monitoring and troubleshooting, service management, and network troubleshooting tools.
- Preferred: Experience in Data Center operations or hardware maintenance, supporting AI/HPC infrastructure, familiarity with monitoring and orchestration tools.
Benefits:
- A culture that values authenticity and diversity of thoughts and backgrounds.
- An inclusive and respectable environment with open workspaces and exciting start-up spirit.
- Fast-growing company with networking opportunities.
- Ability to contribute directly and make an impact on the future of the digital asset industry.
- Attractive welfare benefits and developmental opportunities such as training and mentoring.