jiriOndrusek-agent commented on code in PR #9290:
URL: https://github.com/apache/camel-quarkus/pull/9290#discussion_r4218984504


##########
extensions/langchain4j-ingest/runtime/src/main/doc/usage.adoc:
##########
@@ -115,6 +136,26 @@ A missing parser extension fails the build naming the 
artifact, the same way a m
 The Java twin is `IngestPipeline.parser("tika")`, and a consumer-fed pipeline 
parses the same way — the payload the consumer delivers goes to the parser 
first.
 With a parser, `max-document-size` also rejects a raw payload larger than the 
cap (in bytes) before it reaches the parser, on top of capping the extracted 
text (in characters).
 
+=== Ingesting media: images, audio, video and PDF
+
+`modality=media` embeds each document whole, as one vector, instead of reading 
it as text: the bytes go to a multimodal embedding model as audio, an image, 
video or a PDF, told apart by the MIME type, and the vector is stored with a 
placeholder segment whose text is the document id, stamped with the usual 
identity metadata, so retrieval cites it like any other.
+The model must declare the matching content type in its 
`supportedContentTypes()`: the pipeline fails to start with a text-only model, 
and a document whose medium the model does not declare fails before its bytes 
are read.
+`parser` and `document-splitter` must not be set, the splitter sizes and 
`embedding-batch-size` do not apply, and `max-document-size` and 
`filter.min-document-size` count bytes.
+
+[source,properties]
+----
+quarkus.camel.langchain4j.ingest.photos.source.directory=/var/data/photos
+quarkus.camel.langchain4j.ingest.photos.modality=media
+quarkus.camel.langchain4j.ingest.photos.embedding-model=my-multimodal-model
+quarkus.camel.langchain4j.ingest.photos.embedding-store=photos
+quarkus.camel.langchain4j.ingest.photos.filter.include-id=**/*.png,**/*.jpg
+----
+
+The MIME type handed to the model comes from `content-type` or, when unset, 
from the file extension through Camel's MIME table (wav, mp3, flac, ogg, opus, 
m4a, aac and aiff for audio; png, jpg, gif and webp for images; mp4, mov and 
webm for video; pdf); a file whose type cannot be determined fails the 
exchange, and `content-type` without `modality=media` is rejected like the 
parser and splitter conflicts.
+On a directory pipeline such a file — `notes.txt`, `Thumbs.db` — fails again 
on every poll, so restrict the ids with `filter.include-id`, as above: it is 
checked before the media type.

Review Comment:
   True: the rollback releases the file's key. I'd keep it documented for now:
   - The extension can't tell this failure from the others without copying the 
component's MIME table (`MediaTypes` is package-private).
   - An oversized file is retried the same way on purpose 
(`Langchain4jIngestTikaTest`).
   - With a persistent register, a committed key would keep the file skipped 
even after a config fix.
   
   Answering `filtered` for an untypable id fits better in the component. I can 
open a Camel issue.
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to