Zhuoxi2000 opened a new pull request, #1164:
URL: https://github.com/apache/flink-agents/pull/1164

   <!--
   * Thank you very much for contributing to Flink Agents.
   * Please add the relevant components in the PR title. E.g., [api], 
[runtime], [java], [python], [hotfix], etc.
   -->
   
   <!-- Please link the PR to the relevant issue(s). Hotfix doesn't need this. 
-->
   Linked issue: #1059
   
   ### Purpose of change
   
   <!-- What is the purpose of this change? -->
   
   A user `ChatMessage` with image, audio or document blocks now reaches 
OpenAI, Azure OpenAI and vLLM with its media, in Java and Python. Until now 
every provider sent only the text, silently dropping media. This is the first 
provider step of #1059 Phase 2; Ollama and explicit errors for the remaining 
providers follow in separate PRs.
   
   #### Runtime flow
   
   1. The three connections share one Chat Completions converter per language: 
`OpenAIChatCompletionsUtils.convertToOpenAIMessage` (Java) and 
`convert_to_openai_message` (Python).
   2. A system, assistant or tool message is first checked for media blocks and 
rejected if it has any.
   3. A user message without media is sent with its text projection as a 
string, as before. With media, every block is mapped in order to a content 
part, and the first block without a part throws.
   
   #### Key decisions
   
   * Text-only user messages keep string content, so existing requests and 
OpenAI-compatible servers see no change.
   * Unsupported blocks throw a new `UnsupportedContentBlockException` / 
`UnsupportedContentBlockError` (an `IllegalArgumentException` / `ValueError`) 
from the api module, rather than being dropped or converted; the later provider 
PRs reuse it. Its message names the block type, media type and source type, 
never the payload or URL.
   * Base64 images and documents are sent as `data:` URIs, the form OpenAI 
documents for `image_url` and `file_data`.
   * Documents are not restricted to PDF: OpenAI rejects other types itself; 
compatible servers may accept more.
   * Video, audio by URL and documents by URL are rejected: Chat Completions 
has no part for them. vLLM's `video_url` extension is left out.
   
   ### Behavioral Semantics
   
   <!-- For a non-trivial code change whose implementation is largely 
AI-assisted: interaction decisions, behavioral contracts, and failure behavior. 
See `contribution-guides/ai-assisted-pr.md`. Remove this heading and this 
comment otherwise. -->
   
   #### Interaction decisions
   
   | Role | Blocks | Result |
   |---|---|---|
   | user | text only, or none | `content` is the text projection string |
   | user | media, every block mappable | `content` is a part list in block 
order, text blocks included |
   | user | a block without a part | `UnsupportedContentBlockException`; no 
request is sent |
   | system / assistant / tool | text only | unchanged |
   | system / assistant / tool | any media | 
`UnsupportedContentBlockException`; no request is sent |
   
   #### Behavioral contracts
   
   1. A user message without media is sent with string content equal to its 
text projection.
   2. A user message with media is sent as content parts in block order, one 
`text` part per `TextBlock`.
   3. `ImageBlock` becomes `image_url` with the URL, or 
`data:<media_type>;base64,<data>`.
   4. `AudioBlock` with Base64 data becomes `input_audio`, format `wav` for 
`audio/wav` (`audio/wave`, `audio/x-wav`, `audio/vnd.wave`) and `mp3` for 
`audio/mpeg` (`audio/mp3`); media-type parameters and case are ignored.
   5. `DocumentBlock` with Base64 data becomes `file` with `file_data` as a 
data URI and `filename` set to the block's `name`, else `document`.
   6. `VideoBlock`, audio or documents by URL, and other audio types throw 
`UnsupportedContentBlockException`.
   7. Media in a system, assistant or tool message throws 
`UnsupportedContentBlockException`.
   8. The exception message never contains the Base64 data or the URL.
   
   Java and Python behave identically for each contract.
   
   #### Failure behavior
   
   * Unsupported block or role: thrown while building the request, so nothing 
is sent; the chat action's error strategy applies as for any provider error. 
The error is deterministic, so RETRY fails the same way on every attempt.
   * A model or server that rejects an accepted part (a non-vision model, a 
non-PDF document on OpenAI): the provider's error propagates unchanged.
   * A tool message with media fails on the media before the existing 
`externalId` check.
   
   ### Tests
   
   <!-- How is this change verified? -->
   
   | Contract | Java: `OpenAIChatCompletionsMultimodalTest` | Python: 
`test_openai_multimodal.py` |
   |---|---|---|
   | 1 | `testTextOnlyUserMessageKeepsStringContent` | 
`test_text_only_user_message_keeps_string_content` |
   | 2, 3 | `testUserMediaBecomesOrderedContentParts` | 
`test_user_media_becomes_ordered_content_parts` |
   | 4 | `testAudioBecomesInputAudio` | `test_audio_becomes_input_audio` |
   | 5 | `testDocumentBecomesFilePart` | `test_document_becomes_file_part` |
   | 6, 8 | `testUnsupportedUserBlocksFailExplicitly` | 
`test_unsupported_user_blocks_fail_explicitly` |
   | 7 | `testMediaOutsideUserMessagesFails` | 
`test_media_outside_user_messages_fails` |
   
   Java tests assert on the request as serialized by the SDK's own mapper, i.e. 
the wire JSON; Python tests assert on the request dicts.
   
   Not verified:
   
   * A live OpenAI, Azure OpenAI or vLLM endpoint: no multimodal request was 
sent to a real model.
   * The audio aliases `audio/wave`, `audio/x-wav`, `audio/vnd.wave` and 
`audio/mp3`, and media-type parameters: mapped in code, not individually tested.
   * `image_url.detail` is never set, so the provider default applies.
   
   <details>
   <summary>Implementation invariants and supporting evidence</summary>
   
   * The converter is shared: Java `OpenAICompletionsConnection`, 
`AzureOpenAIChatModelConnection` and `VLLMChatModelConnection` (a subclass of 
the first); Python `OpenAIChatModelConnection`, 
`AzureOpenAIChatModelConnection` and `VLLMChatModelConnection` likewise.
   * openai-java 4.8.0 takes media parts on user messages only 
(`contentOfArrayOfContentParts`); the system, tool and assistant builders 
accept text parts only, which is why media is rejected by role.
   * `input_audio` formats in openai-java 4.8.0 and openai-python are exactly 
`wav` and `mp3`.
   * OpenAI's file-input guide: Chat Completions accepts `file_data` for PDF 
only, as `data:application/pdf;base64,...`.
   * Verified locally: the Java OpenAI integration module 132/132 and the api 
`ChatMessage*` tests, spotless; Python OpenAI, Azure and vLLM plus chat-message 
tests 191 passed (4 skipped), ruff check and format.
   
   </details>
   
   ### API
   
   <!-- Does this change touches any public APIs? -->
   
   New public types: 
`org.apache.flink.agents.api.chat.messages.UnsupportedContentBlockException` 
(Java) and `flink_agents.api.chat_message.UnsupportedContentBlockError` 
(Python).
   
   Compatibility: requests for text-only messages are unchanged. A user message 
with media used to be sent as its text only; it now carries the media, or fails 
if a block has no Chat Completions part. Media in a non-user message used to be 
dropped; it now fails. Other providers are unchanged.
   
   ### Documentation
   
   <!-- Do not remove this section. Check the proper box only. -->
   
   - [ ] `doc-needed` <!-- Your PR changes impact docs -->
   - [ ] `doc-not-needed` <!-- Your PR changes do not impact docs -->
   - [x] `doc-included` <!-- Your PR already contains the necessary 
documentation updates -->
   
   "Multimodal Input" under OpenAI in `chat_models.md`, linked from Azure 
OpenAI and vLLM.
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   <!-- Do not remove this section. Check the proper box only. -->
   
   - [x] Yes
   - [ ] No
   
   If yes, include a `Generated-by: <tool name and version> (<model name and 
version>)` line, for example `Generated-by: Claude Code 2.1.226 (Claude Opus 
4.6)`, in the commit message so it reaches Git history. Repeat the same line 
here for reviewer visibility. See the [ASF generative tooling 
guidance](https://www.apache.org/legal/generative-tooling.html).
   
   Generated-by: Claude Code 2.1.259 (Claude Opus 5.5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to