wenjin272 opened a new issue, #1216:
URL: https://github.com/apache/flink-agents/issues/1216

   ### Search before asking
   
   - [x] I searched the existing issues. This is a focused follow-up to #1059 
and the typed tool-result model in #1185 / #1056. User-media support such as 
#1204 is related but does not cover returning media from tools.
   
   ### Description
   
   Tools can return screenshots, generated charts, documents, or other media 
that the model needs to inspect before continuing. The typed chat model can 
represent these results as ordered text/media blocks inside a 
`ToolResultBlock`, associated with the original tool call. Provider adapters 
must translate those nested blocks into a supported request format to complete 
the tool-use loop.
   
   The adapters in #1185 currently reject media in tool results. OpenAI Chat 
Completions (including Azure OpenAI and vLLM) and Ollama support some media in 
user messages, but still reject it in tool messages. Other integrations also 
retain explicit rejection until their media conversion is implemented. General 
user-media support must not be treated as proof of tool-result media support.
   
   ### Scope
   
   - Establish a provider/model capability matrix for media in tool results, 
covering OpenAI Responses, Anthropic, Gemini, Bedrock, OpenAI-compatible Chat 
Completions, Ollama, DashScope, and watsonx.ai where applicable.
   - Map supported text/media blocks to provider-native tool-result content, 
preserving tool-call correlation, error status, and content order where the 
protocol supports it. For example, Anthropic supports images and documents 
inside `tool_result.content`.
   - For protocols that cannot embed media in a tool-result item, explicitly 
evaluate whether a correctly associated companion media message is supported. 
Document any limits; do not silently stringify or drop media.
   - Keep Java and Python behavior aligned for integrations available in both 
languages. Preserve explicit, actionable errors for unsupported 
role/media/source combinations.
   - Keep tool execution metadata separate from model-visible content.
   
   MCP result conversion is a related follow-up already tracked by #1059; this 
issue focuses on the model-provider side once a typed tool result exists.
   
   ### Acceptance criteria
   
   - Request-conversion tests cover mixed text/media tool results, ordering, 
call IDs, error results, source encodings, and unsupported combinations.
   - At least one supported provider has an end-to-end test/example covering 
model tool call → tool returns media → model consumes that media and continues.
   - Existing text-only tool loops remain functional, including parallel 
tool-call/result correlation.
   - Documentation states support by provider, role, media type, and source 
encoding, with Java/Python tool examples.
   
   References:
   - [Anthropic tool results with images and 
documents](https://platform.claude.com/docs/en/agents-and-tools/tool-use/handle-tool-calls)
   - [OpenAI Agents SDK: images and files from function 
tools](https://openai.github.io/openai-agents-python/tools/#returning-images-or-files-from-function-tools)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to