Position:
Principal Systems Software Engineer - Observability and Telemetry Platform
Company:
NVIDIA
Compensation:
272,000 USD - 431,250 USD
Location:
US, CA, Santa Clara / US, Remote
Employment type:
Full time
Work Arrangement:
Remote
Short Summary:
The Principal Systems Software Engineer at NVIDIA focuses on designing, building, and maintaining large-scale production systems with high efficiency and availability, utilizing software and systems engineering practices. This role emphasizes automation, performance tuning, and the optimization of production systems.
Responsibilities:
- Design, implement, and support operational and reliability aspects of a large-scale Observability & Telemetry collection platform.
- Engage in and improve the lifecycle of services from inception to deployment and refinement.
- Support services pre-launch through system design consulting and capacity management.
- Maintain live services by monitoring availability, latency, and overall system health.
- Scale systems sustainably through automation and push for changes that improve reliability and velocity.
- Practice sustainable incident response and conduct blameless postmortems.
- Participate in an on-call rotation to support production systems.
Requirement:
- BS degree in Computer Science or a related technical field, or equivalent experience.
- 15+ years of experience with infrastructure automation and distributed systems design.
- 8+ years of experience delivering foundational infrastructure and observability platforms.
- Proficiency in Python, Go, Perl, or Ruby.
- In-depth knowledge of Linux, Networking, and Containers.
Benefits:
Eligible for equity and benefits.