Project description
Provide governed data, enterprise knowledge, document intelligence, knowledge graph, and retrieval capabilities that ground AI agents in trusted evidence
Responsibilities
- Design RAG and enterprise-search architectures.
Build ingestion, transformation, enrichment, document extraction, embedding, and indexing pipelines in Databricks.
Create and maintain Azure AI Search, vector search, and access-controlled retrieval indexes.
Implement Azure Document Intelligence pipelines for extracting structured data from PDFs, forms, engineering documents, reports, and scanned content.
Design and maintain knowledge graph models to represent entities, relationships, metadata, lineage, and domain context for agent reasoning and search enrichment.
Implement metadata, filtering, citations, lineage, data freshness, and security-trimmed access controls.
Measure and optimize retrieval relevance and knowledge quality.
Support structured data products used by agents and deterministic services.
SKILLS
Must have
- Bachelor's degree in Computer Science, Data Engineering, Data Science, Information Systems, or a related field.
5+ years of data engineering, search engineering, analytics engineering, or knowledge-platform experience.
Strong SQL skills and experience building production data pipelines in Databricks.
Experience with data quality, metadata, lineage, and governance practices.
Understanding of structured and unstructured enterprise information retrieval, including document extraction and knowledge graph patterns.
Nice to have
Hands-on Azure AI Search, Databricks, RAG, or vector-search experience.
Experience with Azure Document Intelligence, document processing, chunking, embeddings, ranking, metadata, and citation design.
Hands-on knowledge graph experience using Neo4j, RDF, graph databases, semantic modeling, ontology design, or entity-relationship mapping.
Azure Data Engineer or Databricks certification.