RVC
JobsFor Employers
11 jobs
Sort:
1h 19m ago

AI Orchestration Engineer

Bitdeer·AI and Bitcoin mining infrastructure
📡Remote In-Country
|San Jose, United States of America
distributed_systemsgolanginfrastructure_as_codekubernetes+1
Similar jobs
1h 18m ago

HPC Scheduling Engineer

Bitdeer·AI cloud and Bitcoin mining infrastructure
📡Remote In-Country
|San Jose, United States of America
ansibleauthorizationbashcontainerd+11
1h 20m ago

GPU Systems Engineer

Bitdeer·AI and cloud infrastructure
📡Remote In-Country
|San Jose, United States of America
ansiblecci/cdcuda+3
1h 21m ago

AI Storage Infrastructure Engineer

Bitdeer Technologies Group·AI and Bitcoin mining infrastructure
📡Remote In-Country
|San Jose, United States of America
ansibleci/cdkuberneteslinux+2
1h 21m ago

Kubernetes Engineer

Bitdeer·AI and Bitcoin mining infrastructure
📡Remote In-Country
|San Jose, United States of America
ci/cdcloud_servicesdistributed_systemsgolang+2
2h 1m ago

Cloud Network Engineer

Teradata·Technology
📡Remote In-Country
|Mumbai, India
awsazurebashci/cd+21
2h 7m ago

Staff Software Engineer

Ripple·Fintech and cryptocurrency
🏢Hybrid
|Singapore, Singapore
agileansibleawsazure+14
4h 9m ago

Infrastructure Engineer

Komodo Health·Healthcare technology
📡Remote In-Country
|San Francisco, United States of America
$157k - $220k USD/yr
airflowargocdawsbash+17
4h 12m ago

Compute Systems Architect

Anduril Industries·Defense technology
🏢On-site
|Costa Mesa, United States of America
$146k - $222k USD/yr
computer_vision
5h 9m ago

Development Expert

SAP·Enterprise software and cloud services
🏢On-site
|Berlin, Germany
awsazuredistributed_systemsdocker+10
11h 7m ago

Cloud Engineer

🏢On-site
|St. Louis, United States of America
active_directoryapiauthorizationautomation+5

AI Orchestration Engineer

Bitdeer | AI and Bitcoin mining infrastructure
1 hour 19 minutes ago
Remote In-Country
Full Time
United States of America, San Jose, Austin

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/

Position Overview

We are seeking a Staff AI Scheduling & Orchestration Engineer to lead the workload placement logic that defines our AI-native NeoCloud platform. Standard Kubernetes scheduling is insufficient for the demands of large-scale AI; you will be responsible for eliminating "GPU stranding" and maximizing utilization across our expensive compute fleets. This role is pivotal in building a high-performance scheduling fabric that understands the physical realities of our hardware—from NVLink-connected GPU topologies to InfiniBand interconnects. You will work at the intersection of distributed systems and AI, driving the architectural decisions that enable our platform to handle massive-scale distributed training and inference jobs with industry-leading efficiency.

Key Responsibilities

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively.
  • Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics.
  • Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization.
  • Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
  • Drive the reliability and scalability of the scheduling stack, resolving resource contention and deadlock scenarios in large-scale HPC environments.
  • Mentor junior engineers and conduct design reviews to maintain architectural excellence in our orchestration layer.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
  • Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI).
  • Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments.
  • Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference.
  • Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
  • Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals.
  • Ability to work in a high-velocity engineering environment and translate complex, ambiguous requirements into concrete, scalable engineering solutions.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Required Skills

distributed_systemsgolanginfrastructure_as_codekubernetesterraform

Required Languages

🇬🇧 English

Roles
PythonJavaReactTypeScriptNode.jsGoRustDevOpsData scienceProductDesign
Remote
United StatesUnited KingdomCanadaGermanyPolandSpainNetherlandsPortugal
Work type
Fully remoteRemote in-countryHybridSeniorMid-levelJuniorAll jobs
R© 2026 ReVacancybuild a5c5b8de
AboutContactPrivacyCookiesRefundsTerms & ConditionsFor Employers