Zhuoxi2000 opened a new pull request, #1204:
URL: https://github.com/apache/flink-agents/pull/1204

   <!--
   * Thank you very much for contributing to Flink Agents.
   * Please add the relevant components in the PR title. E.g., [api], 
[runtime], [java], [python], [hotfix], etc.
   -->
   
   <!-- Please link the PR to the relevant issue(s). Hotfix doesn't need this. 
-->
   Linked issue: #1059
   
   ### Purpose of change
   
   <!-- What is the purpose of this change? -->
   
   A user `ChatMessage` with image or document blocks now reaches Anthropic 
with its media, in Java and Python, replacing the rejection #1186 added. This 
is a #1059 follow-up after #1164 (OpenAI Chat Completions) and #1178 (Ollama).
   
   #### Runtime flow
   
   1. System, assistant and tool messages are rejected if they have media, and 
otherwise converted as before.
   2. A user message without media keeps string content. With media, every 
block is mapped in order to an Anthropic content block, and the first block 
without one throws.
   
   #### Key decisions
   
   * Text-only user messages keep string content, so existing requests are 
unchanged.
   * Images go by URL or Base64 in the four types Anthropic accepts (JPEG, PNG, 
GIF, WebP); others fail locally.
   * PDFs go by URL or Base64. `text/plain` Base64 data is decoded and sent as 
Anthropic's plain-text document source; plain text by URL has no source type, 
so it is rejected.
   * A document's `name` becomes the document `title`.
   * Tool results stay text-only. Anthropic's `tool_result` can carry images, 
but no provider sends tool-result media yet; that is a follow-up.
   * This is the media half of #1059's Anthropic follow-up. The Python 
`anthropic_content_blocks` workaround stays: it replays the assistant's 
`tool_use` blocks, which content blocks do not model, so replacing it is a 
separate change.
   * `chat()` rethrows `UnsupportedContentBlockException` unchanged; other 
request-building failures stay wrapped as before (the type issue the #1178 
review caught for Ollama).
   
   ### Behavioral Semantics
   
   <!-- For a non-trivial code change whose implementation is largely 
AI-assisted: interaction decisions, behavioral contracts, and failure behavior. 
See `contribution-guides/ai-assisted-pr.md`. Remove this heading and this 
comment otherwise. -->
   
   #### Interaction decisions
   
   | Role | Blocks | Result |
   |---|---|---|
   | user | text only, or none | string content, as before |
   | user | images and PDF or plain-text documents, with or without text | 
content blocks in block order |
   | user | audio, video, other image or document types, plain text by URL | 
`UnsupportedContentBlockException`; nothing sent |
   | system / assistant / tool | any media | 
`UnsupportedContentBlockException`; nothing sent |
   
   #### Behavioral contracts
   
   1. A user message without media is sent with string content equal to its 
text projection.
   2. A user message with media is sent as content blocks in block order, one 
`text` block per `TextBlock`.
   3. `ImageBlock` in `image/jpeg`, `image/png`, `image/gif` or `image/webp` 
becomes `image` with a `url` or `base64` source; media-type parameters and case 
are ignored.
   4. `DocumentBlock` becomes `document`, titled by its `name`: 
`application/pdf` with a `url` or `base64` source, or `text/plain` Base64 data 
as a `text` source holding the decoded UTF-8 text.
   5. Any other block, or media in a system, assistant or tool message, throws 
`UnsupportedContentBlockException` naming the block but not its data or URL.
   6. Plain-text data that is not valid Base64 throws 
`IllegalArgumentException` / `ValueError` ("could not be decoded"), not the 
unsupported-block error.
   7. These errors reach callers of the public `chat()` with their own types.
   
   Java and Python behave identically for each contract.
   
   #### Failure behavior
   
   * Unsupported media or role, and undecodable plain text: thrown while 
building the request; nothing is sent, and the chat action's error strategy 
applies. These errors are deterministic, so RETRY fails the same way each 
attempt.
   * A request Anthropic rejects (an oversized image, an unreachable URL): its 
error propagates as before.
   
   ### Tests
   
   <!-- How is this change verified? -->
   
   | Contract | Java: `AnthropicMultimodalTest` | Python: 
`test_anthropic_multimodal.py` |
   |---|---|---|
   | 1 | `testTextOnlyUserMessageKeepsStringContent` | 
`test_text_only_user_message_keeps_string_content` |
   | 2, 3 | `testImagesBecomeImageBlocks` | `test_images_become_image_blocks` |
   | 4 | `testDocumentsBecomeDocumentBlocks` | 
`test_documents_become_document_blocks` |
   | 5 | `testUnsupportedBlocksFailExplicitly`, 
`testMediaOutsideUserMessagesFails` | 
`test_unsupported_blocks_fail_explicitly`, 
`test_media_outside_user_messages_fails` |
   | 6 | `testInvalidBase64TextDocumentFails` | 
`test_invalid_base64_text_document_fails` |
   | 7 | `testChatKeepsMediaErrorTypes` | the tests above all call `chat()` |
   
   Java tests assert on the SDK-serialized wire JSON; Python tests on the 
mocked client's request. #1186's Anthropic rejection tests are removed.
   
   Not verified: a live Anthropic endpoint, which needs an API key.
   
   <details>
   <summary>Implementation invariants and supporting evidence</summary>
   
   * anthropic-java: `ImageBlockParam` with `Base64ImageSource` / 
`UrlImageSource`, and `DocumentBlockParam` with `Base64PdfSource`, 
`UrlPdfSource` and `PlainTextSource`; `Base64ImageSource.MediaType` covers 
exactly the four image types.
   * Java decodes plain text with `new String(bytes, UTF_8)` and Python with 
`errors="replace"`, so malformed UTF-8 is handled the same way.
   * Verified locally on main: the Java Anthropic module 107/107 with spotless; 
Python chat-model and chat-message tests 468 passed (18 skipped), with ruff.
   
   </details>
   
   ### API
   
   <!-- Does this change touches any public APIs? -->
   
   No new API; it uses `UnsupportedContentBlockException` / 
`UnsupportedContentBlockError` from #1164.
   
   Compatibility: text-only requests are unchanged. User images and PDF or 
plain-text documents, rejected since #1186, are now sent.
   
   ### Documentation
   
   <!-- Do not remove this section. Check the proper box only. -->
   
   - [ ] `doc-needed` <!-- Your PR changes impact docs -->
   - [ ] `doc-not-needed` <!-- Your PR changes do not impact docs -->
   - [x] `doc-included` <!-- Your PR already contains the necessary 
documentation updates -->
   
   A multimodal note with a block table in the Anthropic section of 
`chat_models.md`; Anthropic moved from the not-yet-supported providers in the 
provider-support note and under Multimodal Input.
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   <!-- Do not remove this section. Check the proper box only. -->
   
   - [x] Yes
   - [ ] No
   
   If yes, include a `Generated-by: <tool name and version> (<model name and 
version>)` line, for example `Generated-by: Claude Code 2.1.226 (Claude Opus 
4.6)`, in the commit message so it reaches Git history. Repeat the same line 
here for reviewer visibility. See the [ASF generative tooling 
guidance](https://www.apache.org/legal/generative-tooling.html).
   
   Generated-by: Claude Code 2.1.259 (Claude Opus 5.5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to