[
https://issues.apache.org/jira/browse/CAMEL-24621?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jiri Ondrusek reassigned CAMEL-24621:
-------------------------------------
Assignee: Jiri Ondrusek (was: Jiří Ondrušek)
> camel-langchain4j-ingest - document ingestion component: the missing
> ingestion half of RAG in Camel
> ---------------------------------------------------------------------------------------------------
>
> Key: CAMEL-24621
> URL: https://issues.apache.org/jira/browse/CAMEL-24621
> Project: Camel
> Issue Type: New Feature
> Reporter: Jiří Ondrušek
> Assignee: Jiri Ondrusek
> Priority: Major
>
> Camel today has strong building blocks for the *retrieval* side of RAG —
> camel-langchain4j-chat and camel-langchain4j-agent consume a vector store,
> camel-langchain4j-embeddingstore searches one — but nothing that packages
> getting documents *into* the store. A user who wants a knowledge base has to
> hand-orchestrate splitting, embedding, batching, metadata and deduplication
> in every route. [design/langchain4j-evolution.adoc, "Phase 4: The RAG and
> AIService
> gap"|https://github.com/apache/camel/blob/main/design/langchain4j-evolution.adoc#phase-4-the-rag-and-aiservice-gap]
> names this gap explicitly ("there was still no way to do RAG end-to-end").
> This ticket proposes closing it with a new component in components/camel-ai.
> *What it brings*
> One line turns any of Camel's 300+ consumers into a RAG ingestion source:
> {code}
> from("aws2-s3://product-docs?deleteAfterRead=false")
> .to("langchain4j-ingest:products?documentIdHeader=CamelAwsS3Key");
> {code}
> The producer splits the payload into overlapping segments (recursive
> splitter, configurable sizes), embeds them in batches of 32 (staying under
> embedding providers' per-request limits), and writes them to any LangChain4j
> EmbeddingStore — Qdrant, Milvus, pgvector, Weaviate, Neo4j and the rest —
> stamping every segment with camel_ingest_pipeline and
> camel_ingest_document_id metadata so retrieval can cite its sources and a
> future synchronising engine can find a document's vectors again. Data is
> written the way LangChain4j itself writes it (the same compatibility argument
> that shaped camel-langchain4j-embeddingstore), so existing retrieval — Camel
> or plain LangChain4j — finds it.
> On top of the endpoint sit two thin layers, so every audience gets its
> natural entry point: a declarative Java API (IngestPipelineDefinition +
> IngestPipelineRouteBuilder — "watch this directory, parse with Tika, dedupe,
> store there" without writing route internals), and kamelets
> (langchain4j-ingest-sink, tika/docling variants, a file source) giving
> camel-jbang, camel-k and Pipe users a zero-route-code experience.
> *Why users will like it*
> * *Nothing new to learn*: it is an ordinary producer endpoint. Store and
> model beans autowire when the registry holds exactly one of each; with zero
> or several candidates the endpoint fails fast with the fix spelled out in the
> message ("name the one to use with embeddingStore=#bean:...").
> * *Continuous by nature*: unlike one-shot loaders (e.g. Easy RAG's
> startup-only file loading), Camel sources poll and stream — the knowledge
> base stays current as S3 buckets, Kafka topics or folders change.
> * *Deduplication out of the box*: an optional IdempotentRepository makes
> redeliveries and re-listings answer "skipped" instead of duplicating vectors
> — first write wins per document id, a blank or failed delivery releases its
> claim, and a shared JDBC register gives correct behaviour across replicas.
> * *Safe defaults*: a knowledge base reads its source, it never consumes it
> (noop, no deletes), half-copied files are waited for, and the document id is
> captured before any parser runs, so a crafted document cannot forge its own
> identity.
> * *Honest result contract*: the reply body is an IngestResult (ingested /
> empty / skipped + segments written), so request-reply callers can react.
> *Provenance and proof*
> The design is not speculative: it is the field-proven engine of the
> camel-quarkus-langchain4j-ingest extension (shipped in Camel Quarkus 3.39.0;
> apache/camel-quarkus PRs #9018, #9040, #9078), ported to plain Camel with
> runtime-neutral naming. Once released, the Quarkus extension will delegate to
> this component, so both runtimes share a single implementation and Camel
> Quarkus keeps only its build-time developer experience — one engine,
> maintained once. A working PoC exists with a full unit suite plus an
> integration test that ingests real files through the declarative API into a
> real Qdrant container with a real all-MiniLM-L6-v2 model and verifies that
> semantic questions sharing no vocabulary with the documents retrieve the
> right ones by meaning.
> *Scope and follow-ups*
> Proposed as Preview. Deliberately append-only in this increment — an edited
> document re-ingests alongside its old segments; replace/delete arrive with a
> later synchronising engine, for which the document-identity metadata laid
> down here is the foundation. Follow-ups tracked separately: contributing the
> kamelets to apache/camel-kamelets, a langChain4jRecursiveTokenizer for the
> Tokenizer SPI, and the Camel Quarkus delegation.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)