jiriOndrusek-agent commented on code in PR #9290:
URL: https://github.com/apache/camel-quarkus/pull/9290#discussion_r4218984504
##########
extensions/langchain4j-ingest/runtime/src/main/doc/usage.adoc:
##########
@@ -115,6 +136,26 @@ A missing parser extension fails the build naming the
artifact, the same way a m
The Java twin is `IngestPipeline.parser("tika")`, and a consumer-fed pipeline
parses the same way — the payload the consumer delivers goes to the parser
first.
With a parser, `max-document-size` also rejects a raw payload larger than the
cap (in bytes) before it reaches the parser, on top of capping the extracted
text (in characters).
+=== Ingesting media: images, audio, video and PDF
+
+`modality=media` embeds each document whole, as one vector, instead of reading
it as text: the bytes go to a multimodal embedding model as audio, an image,
video or a PDF, told apart by the MIME type, and the vector is stored with a
placeholder segment whose text is the document id, stamped with the usual
identity metadata, so retrieval cites it like any other.
+The model must declare the matching content type in its
`supportedContentTypes()`: the pipeline fails to start with a text-only model,
and a document whose medium the model does not declare fails before its bytes
are read.
+`parser` and `document-splitter` must not be set, the splitter sizes and
`embedding-batch-size` do not apply, and `max-document-size` and
`filter.min-document-size` count bytes.
+
+[source,properties]
+----
+quarkus.camel.langchain4j.ingest.photos.source.directory=/var/data/photos
+quarkus.camel.langchain4j.ingest.photos.modality=media
+quarkus.camel.langchain4j.ingest.photos.embedding-model=my-multimodal-model
+quarkus.camel.langchain4j.ingest.photos.embedding-store=photos
+quarkus.camel.langchain4j.ingest.photos.filter.include-id=**/*.png,**/*.jpg
+----
+
+The MIME type handed to the model comes from `content-type` or, when unset,
from the file extension through Camel's MIME table (wav, mp3, flac, ogg, opus,
m4a, aac and aiff for audio; png, jpg, gif and webp for images; mp4, mov and
webm for video; pdf); a file whose type cannot be determined fails the
exchange, and `content-type` without `modality=media` is rejected like the
parser and splitter conflicts.
+On a directory pipeline such a file — `notes.txt`, `Thumbs.db` — fails again
on every poll, so restrict the ids with `filter.include-id`, as above: it is
checked before the media type.
Review Comment:
True: the rollback releases the file's key. I'd keep it documented for now:
- The extension can't tell this failure from the others without copying the
component's MIME table (`MediaTypes` is package-private).
- An oversized file is retried the same way on purpose
(`Langchain4jIngestTikaTest`).
- With a persistent register, a committed key would keep the file skipped
even after a config fix.
Answering `filtered` for an untypable id fits better in the component. I can
open a Camel issue.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]