AI Engineer
This is us
Kaltura’s (NYSE:KLTR) mission is to power any video experience for any organization – live, on-demand, or real-time. We not only want to make using video simpler, but we also want to better people’s lives through video. Founded in 2006, Kaltura is now a global leader in the video market with millions of people using our products daily to teach, learn, watch, connect, and collaborate. Among our customers, you’ll find more than 1000 global, well-known organizations.
15+ years since starting the company, we continue to foster a diverse and collaborative work environment where everyone gets a say. Our team is currently 700+ people, and we’re still growing. We have offices in New York, London, Singapore, and Tel Aviv, but our technology is all in the cloud.
Kaltura has a fast-paced environment where initiative is always encouraged. Together with our hybrid work model and flexible state of mind, you get the right conditions for creative juices to flow freely. Thanks to our long line of products, cultivation of rich collaborative culture and care for each Kalturian, you’ll never run out of room to grow and evolve.
If you don't meet 100% of the requirements below - that's okay, nobody's perfect! We believe in hiring people, not just a list of skills. We encourage you to apply if you think this is a role that would make you excited about coming to work every day.
The Role
You will build the shared platform layer that makes Kaltura's agentic AI extensible, measurable, and self-improving. This is the foundation beneath everything: the orchestration engine, the skills registry, the evaluation harness, the guardrails system, the memory service, the integration gateway, and the tracing infrastructure — consumed by both the real-time conversational and offline video generation runtimes.
Your work is what separates "it works for one customer with engineering involvement" from "it works for many customers across many domains without bespoke engineering per customer." You build the machinery once; customer-facing teams and eventually customers themselves extend it.
Scope
You will work across the platform's core systems: agentic orchestration (planning, sub-agents, parallel execution under deadlines), skills registry, evaluation and tracing, guardrails, integration gateway, memory service, and retrieval infrastructure. The shared machinery that both the real-time and offline runtimes consume.
What You Bring
Required
· 4+ years building production backend/platform systems — distributed services, APIs, async processing, and systems that serve multiple consumers. Python primary; experience with high-throughput, low-latency services.
· Deep hands-on experience with LLMs in production — not just using them, but building the infrastructure around them: orchestration, tool calling, retrieval pipelines, context management, prompt chaining, and multi-agent coordination. Proficiency with LangChain, LangGraph, or equivalent orchestration frameworks.
· Experience building platform primitives — registries, gateways, evaluation frameworks, tracing systems, or similar shared infrastructure that other teams build on top of. You understand contracts, versioning, and what it means to ship a platform rather than a feature.
· Demonstrated ability to make build vs. adopt decisions — you have integrated open-source tooling into production systems, understood its boundaries, and built the proprietary layer where needed.
· Experience with RAG systems at depth — indexing strategies, retrieval evaluation, chunking, re-ranking, hybrid search, and the failure modes (context pollution, retrieval misses, conflicting sources). Familiarity with vector databases (Pinecone, Weaviate, Qdrant, or similar) and embedding models.
· Familiarity with real-time system constraints — you understand how latency budgets, deadlines, and streaming shape architecture differently than batch systems.
· Strong systems design skills — you can design a registry, a gateway, a memory service, or an evaluation harness that serves multiple consumers with different needs (internal teams and external customers, real-time and offline paths).
Strongly Preferred
· Experience with evaluation of LLM systems — designing metrics, building test harnesses, measuring behavioral consistency, detecting regressions. Familiarity with DeepEval, Ragas, or similar frameworks.
· Experience with multi-tenant platforms — tenant isolation, scoped configuration, and the engineering discipline required when one customer's data/behavior must never leak to another.
· Background in speech/multimodal AI pipelines — ASR, TTS, avatar rendering, or video generation pipelines. Understanding how these compose with a reasoning layer.
· Experience with observability and tracing at the platform level — structured per-decision traces, not just application logs. Familiarity with Langfuse, Opik, OpenTelemetry, or similar.
· Familiarity with agentic AI frameworks and patterns — CrewAI, AutoGen, Semantic Kernel, MCP protocol, and the failure modes unique to multi-step AI systems (planning loops, tool selection, delegation, parallel execution).
· Experience with guardrails and safety for AI systems — input filtering, output validation, grounding checks, and the latency trade-offs of validation in real-time streaming.
What Success Looks Like
1 month: Deep understanding of the current architecture (Brain, Conversation Manager, real-time and offline pipelines). First contribution shipped — either to evaluation/tracing adoption or registry infrastructure.
3 months: One major platform component owned and in production — registry, evaluation harness, guardrails, or gateway. Integrated with the real-time pipeline and validated under latency constraints. FDE teams consuming your work.
6 months: Multiple foundation components live and serving both runtimes. Evaluation gates enforced on registry entries. Tracing producing actionable diagnostics. Your components are what make new customer onboardings faster rather than bespoke.
Why This Role
You will build the platform layer that turns proven AI technology into an enterprise-grade system. The engine exists — real-time avatar conversations, low-latency speech, offline video generation. What doesn't exist yet is the shared infrastructure that makes it extensible, measurable, and reliable across many customers. That's what you build.
Every component you ship is consumed by multiple teams and multiple customers. Your work compounds.
Required Skills
Required Languages
🇬🇧 English