[
https://issues.apache.org/jira/browse/CAMEL-24656?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112904#comment-18112904
]
Claus Ibsen commented on CAMEL-24656:
-------------------------------------
Merged to main as a series of commits:
* f33abe089d7c core tool set for local providers, lean cache-stable system
prompt, Ollama keep_alive/num_ctx
* 915e4829f647 AI Tools row in F2 Settings (auto/core/full)
* 79dee237aa11 status document as MCP resources
({{camel://status/<pid>[/<section>]}}) and the {{tui_get_status}} tool
* d0091452c389 trimmed the largest tool descriptions; bounded the AI panel
history (tool results capped at 16k chars, older turns compacted to 400 chars,
max 20 turns)
* 1fd5519c6ca7 log tail as an MCP resource ({{camel://log/<pid>[?lines=n]}})
Final measured static prefix per request (system prompt + tool schemas):
||Provider||Before||After||
|Local (Ollama / localhost), core set of 19 tools|~8.1k tokens|~3.0k tokens|
|Hosted, all 47 tools|~8.1k tokens|~6.9k tokens|
End-to-end with qwen3.6:35b-a3b on an M4 Pro using the real payload: first
question ~5s wall (core) / ~14s (full), follow-up questions ~2s, and switching
the monitored integration no longer invalidates the prompt cache.
_Claude Code on behalf of davsclaus_
> camel-jbang TUI: reduce AI panel (F8) prompt size and make it cache friendly
> for local Ollama models
> ----------------------------------------------------------------------------------------------------
>
> Key: CAMEL-24656
> URL: https://issues.apache.org/jira/browse/CAMEL-24656
> Project: Camel
> Issue Type: Improvement
> Components: camel-jbang, camel-jbang-mcp
> Reporter: Claus Ibsen
> Assignee: Claus Ibsen
> Priority: Major
> Fix For: 4.23.0
>
>
> The TUI AI panel (F8) sends a large static prefix on *every* LLM request: the
> system prompt from AiPanel.buildSystemPrompt() plus the full JSON tool
> definitions from TuiToolRegistry (46 tui_* tools). Measured by serializing
> the real registry the way LlmClient does:
> ||Part||Size||
> |46 tool schemas|~7.0k tokens|
> |system prompt|~960 tokens, of which ~730 is a prose list of the same tools|
> |total static prefix|~8.1k tokens|
> With a cloud provider this is only cost. With a local Ollama model it
> dominates latency: on an Apple M4 Pro a dense 27B model processes prompts at
> ~110 tok/s, so a trivial question ("what model is this") takes ~60s before
> the first output token, and every tool-call round trip inside the agent loop
> repeats the full prefix. The Ollama MLX engine cannot reuse a cached prefix
> for the Qwen 3.x hybrid architecture (mlx-lm issue 980), and the llama.cpp
> engine only reuses it when the prefix is byte-identical across requests.
> Proposed improvements (they stack):
> # *Core tool set for local providers.* When the provider is Ollama (or an
> OpenAI-compatible localhost URL), only send the ~18 tools needed for Q&A and
> troubleshooting: tui_get_state, tui_get_options, tui_get_table, tui_get_log,
> tui_get_errors, tui_get_diagram, tui_get_topology, tui_get_processor_detail,
> tui_catalog_doc, tui_get_history, tui_get_spans, tui_control,
> tui_send_message, tui_get_files, tui_get_readme, tui_navigate,
> tui_set_log_level, tui_filter. The other 28 (draw, canvas, animate, tape,
> send_keys, sleep, wait_for_idle, themes, run_example, ...) are
> automation/demo tools for external MCP agents. Saves ~4.1k tokens per turn
> (~50%). Add a setting (e.g. camel.tui.ai.tools=core|full) and/or a /tools
> slash command to opt back into the full set. The MCP server keeps exposing
> all tools.
> # *Drop the duplicated prose tool list from the system prompt.* The tool
> schemas already carry the descriptions. Saves ~730 tokens for every provider.
> Also resolve the conflicting guidance: the system prompt says "use
> tui_get_state first" while the tui_get_options description says "call this
> FIRST"; each extra round trip is a full prefix re-evaluation locally.
> # *Make the prefix cache-stable.* The system prompt embeds "The user is
> monitoring: <name> (PID <pid>)" before the tool block, so selecting another
> integration or a restart invalidates the whole cached prefix on llama.cpp.
> Move that line into the user message (or let the model call tui_get_state).
> Also set keep_alive and an explicit num_ctx in the Ollama request options so
> the model and its KV cache are not unloaded between turns (default keep_alive
> is 5 minutes).
> # *Trim verbose tool descriptions/schemas.* tui_draw and tui_draw_shape alone
> are ~820 tokens (the shape/color vocabulary is inlined twice); many
> descriptions repeat "WITHOUT navigating to the X tab" which the system prompt
> already states; parameter-less tools like tui_get_state have 500+ char
> descriptions. A 25-35% trim is realistic. Note this also changes what MCP
> clients see.
> # *Cap conversation history.* messages grow unbounded until /clear, and tool
> results (full log/table dumps) stay verbatim in history and are re-processed
> on every call. Keep the last N turns and/or stub tool-result content from
> earlier turns, and consider a max-chars cap on tool results.
> Measured on an M4 Pro 64GB with a 7k-token prompt (system + user),
> think=false:
> ||Model / engine||1st request prompt eval||same-prefix request||decode||
> |qwen3.8:27b-mlx (MLX)|61s (115 tok/s)|61s (no prefix reuse)|30 tok/s|
> |qwen3.8:27b (GGUF, llama.cpp)|63s (112 tok/s)|5.2s|6 tok/s|
> |qwen3.6:35b-a3b (GGUF, llama.cpp)|9.4s (747 tok/s)|1.5s|52 tok/s|
> So besides the code changes, the TUI docs / doctor should recommend a
> mixture-of-experts model such as qwen3.6:35b-a3b for local use instead of a
> dense 27B/32B model, and note that GGUF (llama.cpp engine) gives prefix
> caching while the MLX engine currently does not for Qwen 3.x hybrid models.
> No tests currently pin the system prompt text or the tool list, so the
> changes need new tests rather than updates (AiPanelTest fake clients ignore
> systemPrompt/tools).
> _Analysis by Claude Code on behalf of davsclaus_
--
This message was sent by Atlassian Jira
(v8.20.10#820010)