[
https://issues.apache.org/jira/browse/CAMEL-24794?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Work on CAMEL-24794 started by Claus Ibsen.
-------------------------------------------
> camel-jbang - TUI Ollama tab showing local model performance (throughput,
> context, host, request log)
> -----------------------------------------------------------------------------------------------------
>
> Key: CAMEL-24794
> URL: https://issues.apache.org/jira/browse/CAMEL-24794
> Project: Camel
> Issue Type: New Feature
> Components: camel-jbang
> Reporter: Claus Ibsen
> Assignee: Claus Ibsen
> Priority: Major
>
> Add an Ollama tab to the Camel TUI (More menu, AI group) that shows how a
> model served by Ollama is performing, inspired by the llm-visuals dashboard
> (https://github.com/DingoOz/llm-visuals) but built on what Ollama actually
> exposes.
> Data sources, all discovered automatically:
> * Ollama API (localhost:11434 by default, or the URL the TUI's LLM client
> resolved, including camel infra run ollama): version, loaded models (/api/ps:
> size, size in GPU memory, allocated context length, unload time), installed
> models (/api/tags), model shape (/api/show: architecture, layers, expert
> count and active experts, max context, quantization, capabilities).
> * Per-request statistics from Ollama responses: prompt_eval_count,
> prompt_eval_cached_count, prompt_eval_duration, eval_count, eval_duration,
> load_duration, total_duration, done_reason. Derived: prefill and decode
> tokens per second, time to first token, cache hit ratio, cold versus warm
> request. Requests made by the TUI AI panel and camel ask feed this directly;
> calls made by Camel routes arrive through the GenAI observability spans
> (tokens, duration, route id).
> * Same-machine bonus: Ollama's runner is a llama-server child process whose
> /slots endpoint reports live state (processing, prompt tokens processed,
> cache tokens, tokens decoded, context size, speculative decoding). Polling it
> gives a live decode rate and context fill for every caller, including routes.
> Plus process CPU and memory of ollama serve and the runner, and GPU
> utilization and memory via ioreg on macOS or nvidia-smi on Linux. These
> panels only show when the Ollama host is this machine and degrade gracefully
> otherwise (remote host or container).
> Screen: title with Ollama version and host; a model line (family, parameters,
> quantization, layers, experts, loaded size, GPU share, context, unload
> countdown); THROUGHPUT panel with decode and prefill tokens per second in
> large numerals, time to first token, load time and a sparkline history;
> CONTEXT panel with a fill bar, cache hit and working/idle state; HOST panel
> with GPU and process figures; REQUESTS log merging TUI and route calls with
> in/out tokens, prefill, decode, ttft and total time. One card per model when
> several are loaded.
> Also expose the same data over the TUI MCP server as tui_get_ollama so AI
> agents can reason about local model performance.
> Out of scope: speculative decoding acceptance rate and expert routing (need
> llama-server --metrics or a llama.cpp patch, neither available through
> Ollama), frontier models over ACP.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)