RVC
JobsFor Employers
9 jobsSort: Relevance
13d 6h ago

AI Observability Engineer

Nebius·General · Cloud Computing
🏢Hybrid
|Netherlands
ci/cdgrafanaprometheuspython+1
2d 19h ago

SRE DevOps Engineer

Lumnix·General · Cloud infrastructure
📡Fully Remote
ansiblebashbgpgrafana+9
3d 18h ago

Staff Site Reliability Engineer

Skydio·General · Aerospace
📡Fully Remote
|Relocation
$240k - $300k USD
awsci/cdgolangkubernetes+2
7d 1h ago

Senior Cloud Architect

LivePerson·General · Conversational AI
📡Fully Remote
$150k - $160k USD
bashgolanggcpinfrastructure_as_code+3
8d 11h ago

Site Reliability Engineer

Yuno·General · Financial Technology
📡Fully Remote
awsdockerevent-driven_architectureinfrastructure_as_code+4
10d 2h ago

Senior Site Reliability Engineer

Filevine·General · Legal Technology
📡Fully Remote
$175k - $195k USD
aiawsbashci/cd+7
11d 2h ago

Director of Engineering

Federato·General · Insurance Technology
📡Fully Remote
$250k - $275k USD
security
11d 8h ago

Senior Site Reliability Engineer

Ping Identity·General · Cybersecurity
📡Fully Remote
automationci/cdfluxgitops+3
13d 5h ago

Senior Site Reliability Engineer

Alpaca·General · Financial Services
📡Fully Remote
gitopsgolangkuberneteslinux+2

AI Observability Engineer

Nebius | General | Cloud Computing
Hybrid
Full Time
Netherlands
Source not independently verified
Confirm the company and recruiter through an official company domain before applying. Never pay money, share one-time codes, sign out of your Apple or Google account, install apps or configuration profiles, or run code from a repository sent by a recruiter.
Position:
AI Observability Engineer

Company:
Nebius

Location:
Europe (Remote)

Employment type:
Full Time

Short Summary:
Nebius is building a high-performance AI cloud platform and is looking for an AI Observability Engineer to manage the observability backbone, ensuring the Azure AI platform is visible, measurable, reliable, and self-service.

Responsibilities:
- Stand up and operate LLM and agent monitoring with Langfuse.
- Capture traces, latency, token usage, cost, quality scores, prompt and model-version analytics, and safety signals.
- Build lightweight internal tooling and exporters in Python.
- Design and maintain Grafana dashboards, Prometheus metrics, and the Azure observability stack.
- Instrument platform and AI workloads for health, usage, cost, and SLA reporting.
- Feed telemetry and operational insights into the Platform Engineering backlog.
- Own Terraform IaC and CI/CD for observability tooling.
- Support incident investigation and root-cause analysis.

Requirement:
- 5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering.
- Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana.
- Hands-on experience with Langfuse, Grafana, and Prometheus.
- Experience with Terraform and CI/CD.
- Python skills for instrumentation, exporters, and automation.
- Familiarity with ML workloads and AI-specific metrics.
- Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs.
- Intermediate or higher English.

Benefits:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams

Required Skills

ci/cdgrafanaprometheuspythonterraform

Required Languages

🇬🇧 English

Key competency: General
Roles
PythonJavaReactTypeScriptNode.jsGoRustDevOpsData scienceProductDesign
Remote
United StatesUnited KingdomCanadaGermanyPolandSpainNetherlandsPortugal
Work type
Fully remoteRemote in-countryHybridSeniorMid-levelJuniorAll jobs
R© 2026 ReVacancybuild 584ca0bb
AboutContactPrivacyCookiesRefundsTerms & ConditionsFor Employers