RVC
JobsFor Employers
  1. Jobs
  2. /
  3. GPU Cluster Infrastructure Engineer

GPU Cluster Infrastructure Engineer

Beam | AI inference platform
1h 51m ago
Remote In-Country
Contract
IC
United States of America

Title: GPU Cluster Infrastructure Engineer

Location: New York, NY, US / San Francisco, CA, US / Remote (US)

**Beam** is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

# **About the Role**

We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.

  • Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
  • Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
  • Stand up and validate high-performance storage alongside vendor teams.
  • Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
  • Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.
  • Write runbooks, as-builts, and remote-hands procedures.
  • Provide escalation support after go-live and help our team ramp up.

**Skills & Experience**

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
  • Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.

#

Benefits

  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more

Required Languages

🇬🇧 English

Related searches

  • Remote Jobs in USA
  • Jobs in USA
  • Remote In-Country Jobs
1 jobs
Sort by
1h 51m ago

GPU Cluster Infrastructure Engineer

Beam·AI inference platform
📡Remote In-Country
|United States of America
General Infra/Network
No similar jobs match these filters. Try changing or clearing a filter.
Popular job searches ▸
Remote roles
C# / .NET
Countries
United StatesUnited KingdomCanadaUS & EMEAGermanyPolandSpainNetherlandsPortugal
Experience
SeniorMid-levelJunior
R© 2026 RVC Globalbuild 163c5878
AboutMCPPricingContactPrivacyCookiesRefundsTerms & ConditionsFor Employers