Position:
Senior Software Engineer, AIOps and Observability
Company:
NVIDIA
Compensation:
200,000 USD - 322,000 USD
Location:
US, CA, Santa Clara
Employment type:
Full time
Work Arrangement:
Not specified
Short Summary:
We are looking for a highly skilled Senior Software Engineer to design and develop AIOps & Observability platforms at NVIDIA. The platforms are used by internal teams to monitor, diagnose, and optimize products, assets, and services across various environments.
Responsibilities:
- Lead the design, development, and deployment of AIOps & Observability platforms.
- Drive the technical vision and roadmap for AIOps and Observability initiatives.
- Collaborate with teams to understand observability needs and provide solutions.
- Establish and implement observability standards and processes.
- Research and adopt new observability technologies.
- Provide peer reviews and feedback on performance and security.
- Work with Data scientists to implement machine learning models.
- Handle large volumes of data ensuring quality and compliance.
- Develop scalable, reliable, and distributed systems.
- Build AI agents and observability tools for faster issue resolution.
Requirement:
- Bachelor’s degree in computer science or related field, or equivalent experience.
- 12+ years of experience in product development and full stack engineering.
- 5+ years in developing observability platforms, preferably in cloud-native environments.
- Strong knowledge of observability tools (e.g., Prometheus, Grafana).
- Hands-on knowledge in AIOps tools (e.g., BigPanda, Datadog).
- Experience with Kubernetes, Docker, and microservices architectures.
- Proficient in programming languages such as Go, Python, Java, C#.
- Passionate about observability and high-quality internal platforms.
Benefits:
- Eligible for equity and benefits.
- NVIDIA fosters an inclusive work environment and is an equal opportunity employer.