RVC
JobsFor Employers
9 jobsSort: Relevance
4d 17h ago

GPU Network SRE Expert

Bitdeer·General · Technology
📡Remote In-Country
|San Jose, United States of America
qos
2d 22h ago

SRE DevOps Engineer

Lumnix·General · Cloud infrastructure
📡Fully Remote
ansiblebashbgpgrafana+9
3d 21h ago

Staff Site Reliability Engineer

Skydio·General · Aerospace
📡Fully Remote
|Relocation
$240k - $300k USD
awsci/cdgolangkubernetes+2
7d 4h ago

Senior Cloud Architect

LivePerson·General · Conversational AI
📡Fully Remote
$150k - $160k USD
bashgolanggcpinfrastructure_as_code+3
8d 14h ago

Site Reliability Engineer

Yuno·General · Financial Technology
📡Fully Remote
awsdockerevent-driven_architectureinfrastructure_as_code+4
10d 5h ago

Senior Site Reliability Engineer

Filevine·General · Legal Technology
📡Fully Remote
$175k - $195k USD
aiawsbashci/cd+7
11d 5h ago

Director of Engineering

Federato·General · Insurance Technology
📡Fully Remote
$250k - $275k USD
security
11d 11h ago

Senior Site Reliability Engineer

Ping Identity·General · Cybersecurity
📡Fully Remote
automationci/cdfluxgitops+3
13d 8h ago

Senior Site Reliability Engineer

Alpaca·General · Financial Services
📡Fully Remote
gitopsgolangkuberneteslinux+2

GPU Network SRE Expert

Bitdeer | General | Technology
Remote In-Country
Full Time
United States of America, San Jose, Austin

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

About the Role

You keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.

Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.

What you'll own

  • InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
  • RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
  • UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
  • IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
  • NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
  • Firmware lifecycle across IB switches and HCAs.
  • Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
  • Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.

Feed the AIOps substrate

  • Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
  • Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
  • Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.

Job Requirement:

  • 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
  • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
  • Strong understanding of IB subnet management, partitioning, and QoS
  • Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
  • Proficiency with UFM or equivalent IB fabric management tools
  • Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
  • Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
  • Understanding of NCCL and how GPU communication maps to network topology
  • Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
  • Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Required Skills

qos

Required Languages

🇬🇧 English

Key competency: General
Roles
PythonJavaReactTypeScriptNode.jsGoRustDevOpsData scienceProductDesign
Remote
United StatesUnited KingdomCanadaGermanyPolandSpainNetherlandsPortugal
Work type
Fully remoteRemote in-countryHybridSeniorMid-levelJuniorAll jobs
R© 2026 ReVacancybuild ea4652b0
AboutContactPrivacyCookiesRefundsTerms & ConditionsFor Employers