wenjin272 opened a new issue, #1216: URL: https://github.com/apache/flink-agents/issues/1216
### Search before asking - [x] I searched the existing issues. This is a focused follow-up to #1059 and the typed tool-result model in #1185 / #1056. User-media support such as #1204 is related but does not cover returning media from tools. ### Description Tools can return screenshots, generated charts, documents, or other media that the model needs to inspect before continuing. The typed chat model can represent these results as ordered text/media blocks inside a `ToolResultBlock`, associated with the original tool call. Provider adapters must translate those nested blocks into a supported request format to complete the tool-use loop. The adapters in #1185 currently reject media in tool results. OpenAI Chat Completions (including Azure OpenAI and vLLM) and Ollama support some media in user messages, but still reject it in tool messages. Other integrations also retain explicit rejection until their media conversion is implemented. General user-media support must not be treated as proof of tool-result media support. ### Scope - Establish a provider/model capability matrix for media in tool results, covering OpenAI Responses, Anthropic, Gemini, Bedrock, OpenAI-compatible Chat Completions, Ollama, DashScope, and watsonx.ai where applicable. - Map supported text/media blocks to provider-native tool-result content, preserving tool-call correlation, error status, and content order where the protocol supports it. For example, Anthropic supports images and documents inside `tool_result.content`. - For protocols that cannot embed media in a tool-result item, explicitly evaluate whether a correctly associated companion media message is supported. Document any limits; do not silently stringify or drop media. - Keep Java and Python behavior aligned for integrations available in both languages. Preserve explicit, actionable errors for unsupported role/media/source combinations. - Keep tool execution metadata separate from model-visible content. MCP result conversion is a related follow-up already tracked by #1059; this issue focuses on the model-provider side once a typed tool result exists. ### Acceptance criteria - Request-conversion tests cover mixed text/media tool results, ordering, call IDs, error results, source encodings, and unsupported combinations. - At least one supported provider has an end-to-end test/example covering model tool call → tool returns media → model consumes that media and continues. - Existing text-only tool loops remain functional, including parallel tool-call/result correlation. - Documentation states support by provider, role, media type, and source encoding, with Java/Python tool examples. References: - [Anthropic tool results with images and documents](https://platform.claude.com/docs/en/agents-and-tools/tool-use/handle-tool-calls) - [OpenAI Agents SDK: images and files from function tools](https://openai.github.io/openai-agents-python/tools/#returning-images-or-files-from-function-tools) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
