GPU Cluster Infrastructure Engineer
Title: GPU Cluster Infrastructure Engineer
Location: New York, NY, US / San Francisco, CA, US / Remote (US)
**Beam** is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.
# **About the Role**
We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.
- Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
- Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
- Stand up and validate high-performance storage alongside vendor teams.
- Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
- Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.
- Write runbooks, as-builts, and remote-hands procedures.
- Provide escalation support after go-live and help our team ramp up.
**Skills & Experience**
- You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
- Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
- GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
- Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
- Equally effective on the data center floor and remotely, including directing colo remote hands.
- You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
- Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.
#
Benefits
- Competitive salary and meaningful equity
- Join a fast-growing pre-series A company at the ground floor
- Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
- Opportunities to participate in events across the cloud native community
- Fitness stipend, learning budget, and much, much more
Required Languages
🇬🇧 English