Position:
AI Observability Engineer
Company:
Nebius
Location:
Europe (Remote)
Employment type:
Full Time
Short Summary:
Nebius is building a high-performance AI cloud platform and is looking for an AI Observability Engineer to manage the observability backbone, ensuring the Azure AI platform is visible, measurable, reliable, and self-service.
Responsibilities:
- Stand up and operate LLM and agent monitoring with Langfuse.
- Capture traces, latency, token usage, cost, quality scores, prompt and model-version analytics, and safety signals.
- Build lightweight internal tooling and exporters in Python.
- Design and maintain Grafana dashboards, Prometheus metrics, and the Azure observability stack.
- Instrument platform and AI workloads for health, usage, cost, and SLA reporting.
- Feed telemetry and operational insights into the Platform Engineering backlog.
- Own Terraform IaC and CI/CD for observability tooling.
- Support incident investigation and root-cause analysis.
Requirement:
- 5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering.
- Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana.
- Hands-on experience with Langfuse, Grafana, and Prometheus.
- Experience with Terraform and CI/CD.
- Python skills for instrumentation, exporters, and automation.
- Familiarity with ML workloads and AI-specific metrics.
- Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs.
- Intermediate or higher English.
Benefits:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams