Zhuoxi2000 opened a new pull request, #1204:
URL: https://github.com/apache/flink-agents/pull/1204
<!--
* Thank you very much for contributing to Flink Agents.
* Please add the relevant components in the PR title. E.g., [api],
[runtime], [java], [python], [hotfix], etc.
-->
<!-- Please link the PR to the relevant issue(s). Hotfix doesn't need this.
-->
Linked issue: #1059
### Purpose of change
<!-- What is the purpose of this change? -->
A user `ChatMessage` with image or document blocks now reaches Anthropic
with its media, in Java and Python, replacing the rejection #1186 added. This
is a #1059 follow-up after #1164 (OpenAI Chat Completions) and #1178 (Ollama).
#### Runtime flow
1. System, assistant and tool messages are rejected if they have media, and
otherwise converted as before.
2. A user message without media keeps string content. With media, every
block is mapped in order to an Anthropic content block, and the first block
without one throws.
#### Key decisions
* Text-only user messages keep string content, so existing requests are
unchanged.
* Images go by URL or Base64 in the four types Anthropic accepts (JPEG, PNG,
GIF, WebP); others fail locally.
* PDFs go by URL or Base64. `text/plain` Base64 data is decoded and sent as
Anthropic's plain-text document source; plain text by URL has no source type,
so it is rejected.
* A document's `name` becomes the document `title`.
* Tool results stay text-only. Anthropic's `tool_result` can carry images,
but no provider sends tool-result media yet; that is a follow-up.
* This is the media half of #1059's Anthropic follow-up. The Python
`anthropic_content_blocks` workaround stays: it replays the assistant's
`tool_use` blocks, which content blocks do not model, so replacing it is a
separate change.
* `chat()` rethrows `UnsupportedContentBlockException` unchanged; other
request-building failures stay wrapped as before (the type issue the #1178
review caught for Ollama).
### Behavioral Semantics
<!-- For a non-trivial code change whose implementation is largely
AI-assisted: interaction decisions, behavioral contracts, and failure behavior.
See `contribution-guides/ai-assisted-pr.md`. Remove this heading and this
comment otherwise. -->
#### Interaction decisions
| Role | Blocks | Result |
|---|---|---|
| user | text only, or none | string content, as before |
| user | images and PDF or plain-text documents, with or without text |
content blocks in block order |
| user | audio, video, other image or document types, plain text by URL |
`UnsupportedContentBlockException`; nothing sent |
| system / assistant / tool | any media |
`UnsupportedContentBlockException`; nothing sent |
#### Behavioral contracts
1. A user message without media is sent with string content equal to its
text projection.
2. A user message with media is sent as content blocks in block order, one
`text` block per `TextBlock`.
3. `ImageBlock` in `image/jpeg`, `image/png`, `image/gif` or `image/webp`
becomes `image` with a `url` or `base64` source; media-type parameters and case
are ignored.
4. `DocumentBlock` becomes `document`, titled by its `name`:
`application/pdf` with a `url` or `base64` source, or `text/plain` Base64 data
as a `text` source holding the decoded UTF-8 text.
5. Any other block, or media in a system, assistant or tool message, throws
`UnsupportedContentBlockException` naming the block but not its data or URL.
6. Plain-text data that is not valid Base64 throws
`IllegalArgumentException` / `ValueError` ("could not be decoded"), not the
unsupported-block error.
7. These errors reach callers of the public `chat()` with their own types.
Java and Python behave identically for each contract.
#### Failure behavior
* Unsupported media or role, and undecodable plain text: thrown while
building the request; nothing is sent, and the chat action's error strategy
applies. These errors are deterministic, so RETRY fails the same way each
attempt.
* A request Anthropic rejects (an oversized image, an unreachable URL): its
error propagates as before.
### Tests
<!-- How is this change verified? -->
| Contract | Java: `AnthropicMultimodalTest` | Python:
`test_anthropic_multimodal.py` |
|---|---|---|
| 1 | `testTextOnlyUserMessageKeepsStringContent` |
`test_text_only_user_message_keeps_string_content` |
| 2, 3 | `testImagesBecomeImageBlocks` | `test_images_become_image_blocks` |
| 4 | `testDocumentsBecomeDocumentBlocks` |
`test_documents_become_document_blocks` |
| 5 | `testUnsupportedBlocksFailExplicitly`,
`testMediaOutsideUserMessagesFails` |
`test_unsupported_blocks_fail_explicitly`,
`test_media_outside_user_messages_fails` |
| 6 | `testInvalidBase64TextDocumentFails` |
`test_invalid_base64_text_document_fails` |
| 7 | `testChatKeepsMediaErrorTypes` | the tests above all call `chat()` |
Java tests assert on the SDK-serialized wire JSON; Python tests on the
mocked client's request. #1186's Anthropic rejection tests are removed.
Not verified: a live Anthropic endpoint, which needs an API key.
<details>
<summary>Implementation invariants and supporting evidence</summary>
* anthropic-java: `ImageBlockParam` with `Base64ImageSource` /
`UrlImageSource`, and `DocumentBlockParam` with `Base64PdfSource`,
`UrlPdfSource` and `PlainTextSource`; `Base64ImageSource.MediaType` covers
exactly the four image types.
* Java decodes plain text with `new String(bytes, UTF_8)` and Python with
`errors="replace"`, so malformed UTF-8 is handled the same way.
* Verified locally on main: the Java Anthropic module 107/107 with spotless;
Python chat-model and chat-message tests 468 passed (18 skipped), with ruff.
</details>
### API
<!-- Does this change touches any public APIs? -->
No new API; it uses `UnsupportedContentBlockException` /
`UnsupportedContentBlockError` from #1164.
Compatibility: text-only requests are unchanged. User images and PDF or
plain-text documents, rejected since #1186, are now sent.
### Documentation
<!-- Do not remove this section. Check the proper box only. -->
- [ ] `doc-needed` <!-- Your PR changes impact docs -->
- [ ] `doc-not-needed` <!-- Your PR changes do not impact docs -->
- [x] `doc-included` <!-- Your PR already contains the necessary
documentation updates -->
A multimodal note with a block table in the Anthropic section of
`chat_models.md`; Anthropic moved from the not-yet-supported providers in the
provider-support note and under Multimodal Input.
### Was this patch authored or co-authored using generative AI tooling?
<!-- Do not remove this section. Check the proper box only. -->
- [x] Yes
- [ ] No
If yes, include a `Generated-by: <tool name and version> (<model name and
version>)` line, for example `Generated-by: Claude Code 2.1.226 (Claude Opus
4.6)`, in the commit message so it reaches Git history. Repeat the same line
here for reviewer visibility. See the [ASF generative tooling
guidance](https://www.apache.org/legal/generative-tooling.html).
Generated-by: Claude Code 2.1.259 (Claude Opus 5.5)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]