Position:
Senior Solutions Architect, AI Factory Deployment - NVIS
Company:
NVIDIA
Compensation:
Base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.
Location:
US, Remote
Employment type:
Full time
Work Arrangement:
Remote
Short Summary:
We are in search of a curious and motivated Senior Solutions Architect to join our NVIDIA Infrastructure Specialists team. You will support the creation, implementation, and verification of AI factories, focusing on running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters.
Responsibilities:
- Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
- Validate configurations against guidelines for NCCL, collectives, and distributed training frameworks.
- Run key AI/LLM benchmarks — setup, orchestration, result collection, and analysis.
- Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
- Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
- Build automation using Python and Shell for conducting benchmarks, retrieving results, and completing regression checks.
- Analyze communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
- Help identify and recommend improvements to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
- Work closely with hardware, software, networking, and product teams to prepare AI factories for customer use.
- Contribute to documentation and readiness materials for internal and customer-facing teams.
Requirement:
- Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or a related field.
- 5+ years of experience managing Linux-based systems in HPC, distributed systems, or AI/ML environments.
- Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, including some exposure to NCCL.
- Practical knowledge of collective communication patterns like AllReduce and AllToAll, and their application in ML/LLM training.
- Skilled in Python and Shell/Bash for scripting, automation, and tooling.
- Strong communication skills and the ability to work effectively with cross-functional teams.
Benefits:
You will also be eligible for equity and benefits.