About Our Client
Our client is a leading enterprise data platform company building an open, high-performance data lakehouse for AI and analytical workloads. The platform combines an intelligent SQL query engine, an AI-ready semantic layer, and an open catalog built on Apache Iceberg β enabling Fortune 500 companies across finance, energy, manufacturing, and logistics to unify, query, and govern data at massive scale across cloud and on-premise sources.
About the Role
We are looking for a Senior Software Engineer to build and operate production AI agent and search services on top of a large-scale data platform. This is hands-on work at the intersection of distributed backend engineering and applied AI: multi-agent workflows, tool-execution infrastructure, retrieval and relevance, and integrations with multiple LLM providers β all running as real, customer-facing production services with the reliability expectations that implies.
Responsibilities
Design and implement scalable multi-agent workflows (including conversation state), tool-execution infrastructure, and data-aware tools working with catalogs, schemas, lineage, SQL, jobs, semantic metadata, data preparation, and visualizations
Build search and retrieval capabilities with OpenSearch or equivalent β indexing, relevance tuning, hybrid keyword + vector search
Integrate multiple LLM/model providers with secure configuration and graceful degradation
Develop APIs and event contracts across REST, streaming/SSE, gRPC, protobuf, and SDKs, including UI integrations
Operate, monitor, and improve production services β service health, logs, metrics, traces, search-cluster health, latency, and error rates; participate in the team's production-support model
Investigate incidents, drive root-cause analysis, and ship durable fixes
Design for permissions, tool approvals, multi-tenancy, and customer-data governance
Maintain dependencies, apply security fixes, and reduce tech debt; produce designs, tests, runbooks, and documentation; lead technical discussions and review code
Required Qualifications
B.
S. or M.
S. in Computer Science, Computer Engineering, or a related field β or equivalent practical experience
7+ years of relevant experience, or equivalent demonstrated senior-level impact
Strong Java software engineering skills
Strong grasp of concurrency, data structures, API design, testing, maintainability, and backward compatibility
Solid background in distributed systems, microservices, cloud services, or data-intensive backends
Hands-on experience with search technologies β OpenSearch, Elasticsearch, or Lucene β including vector search, embeddings, or relevance systems
Experience building LLM applications, agent/tool workflows, RAG systems, or model-provider integrations
Strong production instincts: diagnosing failures via logs, metrics, and traces; designing for fault tolerance, retries, cancellation, and recovery
SQL and familiarity with data platform concepts: catalogs, metadata, schemas, governance, lineage
Comfort leading ambiguous, cross-service work with clear stakeholder communication and strong written documentation
Comfortable with AI-assisted development workflows β using modern AI tools productively while critically validating their output
English Upper-Intermediate or higher (B2+) β daily written and verbal communication with a US-based engineering team
Availability to work EU business hours shifted 2β3 hours later for daily overlap with US West Coast mornings
Desired Skills
LangChain4j or similar agent frameworks
Python for AI services, evaluation harnesses, or sandboxed execution
React and TypeScript
gRPC, protobuf, Quarkus, Kubernetes, MongoDB, KV stores, event-driven architectures
Apache Arrow, Apache Iceberg, OpenLineage, semantic layers, or data-governance systems
AI quality evaluation: LLM-as-judge, model benchmarking, production AI metrics
OpenSearch operations: index lifecycle management, relevance tuning, incident response
Open-source contributions
Details
Engagement: Long-term contract
Location: Europe (EU / EEA / UK), remote
Working hours: EU business hours, shifted 2β3 hours later for daily overlap with the US West Coast team