[ 
https://issues.apache.org/jira/browse/CAMEL-24621?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jiri Ondrusek reassigned CAMEL-24621:
-------------------------------------

    Assignee: Jiri Ondrusek  (was: Jiří Ondrušek)

> camel-langchain4j-ingest - document ingestion component: the missing 
> ingestion half of RAG in Camel
> ---------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-24621
>                 URL: https://issues.apache.org/jira/browse/CAMEL-24621
>             Project: Camel
>          Issue Type: New Feature
>            Reporter: Jiří Ondrušek
>            Assignee: Jiri Ondrusek
>            Priority: Major
>
> Camel today has strong building blocks for the *retrieval* side of RAG — 
> camel-langchain4j-chat and camel-langchain4j-agent consume a vector store, 
> camel-langchain4j-embeddingstore searches one — but nothing that packages 
> getting documents *into* the store. A user who wants a knowledge base has to 
> hand-orchestrate splitting, embedding, batching, metadata and deduplication 
> in every route. [design/langchain4j-evolution.adoc, "Phase 4: The RAG and 
> AIService 
> gap"|https://github.com/apache/camel/blob/main/design/langchain4j-evolution.adoc#phase-4-the-rag-and-aiservice-gap]
>  names this gap explicitly ("there was still no way to do RAG end-to-end"). 
> This ticket proposes closing it with a new component in components/camel-ai.
> *What it brings*
> One line turns any of Camel's 300+ consumers into a RAG ingestion source:
> {code}
> from("aws2-s3://product-docs?deleteAfterRead=false")
>   .to("langchain4j-ingest:products?documentIdHeader=CamelAwsS3Key");
> {code}
> The producer splits the payload into overlapping segments (recursive 
> splitter, configurable sizes), embeds them in batches of 32 (staying under 
> embedding providers' per-request limits), and writes them to any LangChain4j 
> EmbeddingStore — Qdrant, Milvus, pgvector, Weaviate, Neo4j and the rest — 
> stamping every segment with camel_ingest_pipeline and 
> camel_ingest_document_id metadata so retrieval can cite its sources and a 
> future synchronising engine can find a document's vectors again. Data is 
> written the way LangChain4j itself writes it (the same compatibility argument 
> that shaped camel-langchain4j-embeddingstore), so existing retrieval — Camel 
> or plain LangChain4j — finds it.
> On top of the endpoint sit two thin layers, so every audience gets its 
> natural entry point: a declarative Java API (IngestPipelineDefinition + 
> IngestPipelineRouteBuilder — "watch this directory, parse with Tika, dedupe, 
> store there" without writing route internals), and kamelets 
> (langchain4j-ingest-sink, tika/docling variants, a file source) giving 
> camel-jbang, camel-k and Pipe users a zero-route-code experience.
> *Why users will like it*
> * *Nothing new to learn*: it is an ordinary producer endpoint. Store and 
> model beans autowire when the registry holds exactly one of each; with zero 
> or several candidates the endpoint fails fast with the fix spelled out in the 
> message ("name the one to use with embeddingStore=#bean:...").
> * *Continuous by nature*: unlike one-shot loaders (e.g. Easy RAG's 
> startup-only file loading), Camel sources poll and stream — the knowledge 
> base stays current as S3 buckets, Kafka topics or folders change.
> * *Deduplication out of the box*: an optional IdempotentRepository makes 
> redeliveries and re-listings answer "skipped" instead of duplicating vectors 
> — first write wins per document id, a blank or failed delivery releases its 
> claim, and a shared JDBC register gives correct behaviour across replicas.
> * *Safe defaults*: a knowledge base reads its source, it never consumes it 
> (noop, no deletes), half-copied files are waited for, and the document id is 
> captured before any parser runs, so a crafted document cannot forge its own 
> identity.
> * *Honest result contract*: the reply body is an IngestResult (ingested / 
> empty / skipped + segments written), so request-reply callers can react.
> *Provenance and proof*
> The design is not speculative: it is the field-proven engine of the 
> camel-quarkus-langchain4j-ingest extension (shipped in Camel Quarkus 3.39.0; 
> apache/camel-quarkus PRs #9018, #9040, #9078), ported to plain Camel with 
> runtime-neutral naming. Once released, the Quarkus extension will delegate to 
> this component, so both runtimes share a single implementation and Camel 
> Quarkus keeps only its build-time developer experience — one engine, 
> maintained once. A working PoC exists with a full unit suite plus an 
> integration test that ingests real files through the declarative API into a 
> real Qdrant container with a real all-MiniLM-L6-v2 model and verifies that 
> semantic questions sharing no vocabulary with the documents retrieve the 
> right ones by meaning.
> *Scope and follow-ups*
> Proposed as Preview. Deliberately append-only in this increment — an edited 
> document re-ingests alongside its old segments; replace/delete arrive with a 
> later synchronising engine, for which the document-identity metadata laid 
> down here is the foundation. Follow-ups tracked separately: contributing the 
> kamelets to apache/camel-kamelets, a langChain4jRecursiveTokenizer for the 
> Tokenizer SPI, and the Camel Quarkus delegation.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to