wenjin272 opened a new pull request, #1185: URL: https://github.com/apache/flink-agents/pull/1185
Linked issue: Closes #1056 ### Purpose of change #### Outcome and runtime flow Represent conversation content as ordered typed blocks in Java and Python. Chat models return `ChatResult`, containing an assistant `ChatMessage`, model/response identifiers, token usage, finish reason and result metadata. This separates conversation history from per-invocation output previously mixed into `extra_args` and untyped tool-call maps. `ChatRequestEvent` -> chat action -> provider adapter -> `ChatResult`. The action reads usage and finish reason, stores only the assistant message in history, and dispatches typed tool calls. Tool execution returns `ToolResponse`; history records its model-facing content as `ToolResultBlock` under the same call ID. Structured output and routing information remain on `ChatResponseEvent`. #### Key decisions and review order The three commits group core contracts/execution, provider adaptations, and downstream consumers/bridges/fixtures. Review the first two for design and protocol behavior; apply the complete series for repository buildability. Keep `ToolResponse` separate from `ToolResultBlock`: execution status and timing are distinct from model-facing conversation content. Keep ordinary mutable metadata/input maps without deep-freezing arbitrary values. Stateless provider conversions use local static utilities, with protocol field constants. ### Behavioral Semantics #### Interaction decisions | Input / condition | Behavior | | --- | --- | | SYSTEM / USER / ASSISTANT | SYSTEM accepts text; USER accepts text/media; ASSISTANT accepts text/media/reasoning/tool calls. | | TOOL message | Exactly one ToolResultBlock; its content accepts text/media. Each provider checks its supported media. | | Assistant text plus reasoning/tool calls | Block order survives serialization; text and tool-call properties project only matching top-level blocks. | | Provider reasoning with continuation metadata | Anthropic and Bedrock retain signed/redacted content; Gemini retains thought signatures; OpenAI Responses retains native reasoning items. Ollama stores reasoning but does not replay it. | | Structured output plus chat history | Parsed output is an event attribute; the original assistant message remains in history. | | Java/Python boundary or restored event | Convert the same typed message/result shape, including nested blocks and tool responses. | #### Contracts and failure behavior - Tool calls have explicit IDs, names and input maps. Duplicate call IDs within a message and invalid role/block combinations fail validation. - Tool results reference the same call ID; provider IDs are retained when available and adapters generate IDs when absent. Gemini-generated IDs are not echoed as native IDs. - Usage distinguishes unknown from zero. Metrics read typed usage; finish reasons remain strings, with existing canonical limit/filter mappings used by the execution guard. - Invalid model output, unsupported media and structured-output parsing errors follow validation/action failure paths; service errors retain the existing retry policy and exhausted calls emit failed response events. - Event Log sanitizes media payloads/URLs while retaining reasoning, tool input and metadata; transport/state serialization preserves full content. Mutable nested maps can be shared and must not be treated as deep snapshots. ### Tests | Contract | Coverage | | --- | --- | | Role validation, ordered blocks, result envelope, usage and finish reasons | Java ChatMessageTest/ChatMessageSerializationTest; Python test_chat_message.py/test_chat_result.py | | Original message retained with structured output; typed metrics and tool lifecycle | Java ChatModelActionTest/ToolCallActionTest; corresponding Python action and token metric tests | | Provider mapping and reasoning continuation | Anthropic, Bedrock, Gemini, OpenAI and Ollama provider tests, including serialization-before-replay cases | | Ollama user images survive the new model | Java OllamaMultimodalTest and Python test_ollama_multimodal.py | | Failed events/tools and cross-language wire shape | ChatResponseEventTest, CrossLanguageEventSnapshotTest and Python snapshot tests | | Bridge/state/log consumers | JavaResourceAdapterTest, ActionStateSerdeTest, FileEventLoggerTest and Python runtime conversion tests | After rebasing onto main (`99103da67`): full Java reactor compilation/install succeeded; Java non-E2E tests passed (2627 passed, 44 skipped), excluding `FlussActionStateStoreIntegrationTest` because its embedded service previously blocked during initialization. Python non-integration tests: 1704 passed, 14 skipped. Mock chat MiniCluster E2E: 2 passed. Spotless and Ruff checks for changed files passed. Not verified: live model services, exhaustive parity across providers, and upgrade/recovery from the old message wire format. This PR intentionally breaks that format. ### API Breaking Java/Python change: chat methods return `ChatResult`; replace `extra_args`/`extraArgs` and untyped `tool_calls` with typed blocks and scoped metadata. History still consists of `ChatMessage`; text convenience factories remain available. Use event structured-output accessors for parsed results. Existing serialized messages/events/checkpoints require migration; no compatibility layer is provided. ### Documentation - [ ] `doc-needed` - [ ] `doc-not-needed` - [x] `doc-included` ### Was this patch authored or co-authored using generative AI tooling? - [x] Yes - [ ] No Generated-by: Codex 0.153.4 (GPT-6) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
