Claus Ibsen created CAMEL-25357:
-----------------------------------
Summary: camel-semantic - dev console for semantic evaluations and
experts
Key: CAMEL-25357
URL: https://issues.apache.org/jira/browse/CAMEL-25357
Project: Camel
Issue Type: New Feature
Components: camel-jbang, camel-ai
Reporter: Claus Ibsen
cc [~ldemasi]
h2. Motivation
With experts and capabilities from CAMEL-25310 (and fixed-purpose experts such
as the prompt-injection expert in CAMEL-25311), a running Camel application
makes semantic decisions in its routes. Today nothing shows which evaluations
are declared, which expert each one uses, or what the recent evaluations
returned. camel-semantic has no dev console and emits no observability events.
This issue depends on CAMEL-25310 (PR #27373), since experts and
{{SemanticCapabilities}} come from there.
h2. Proposal
h3. 1. A {{semantic}} dev console in camel-semantic
* *Declarations*: name, result type (BOOLEAN / CHOICE / SCORE), selected
expert, threshold and uncertainty policy, whether instructions are set.
* *Experts*: configured instances, provider and pinned artifact/model identity,
and their {{SemanticCapabilities}} (supported result types, instructions
required/optional/unsupported, fixed meaning such as "positive = prompt
injection detected", meaning of probability/confidence).
* *Recent evaluations*: a bounded ring buffer. Each entry has the time,
route/node, evaluation name, expert, answer/label, probability, confidence,
threshold decision (match / no match / uncertain), latency and error.
** The evaluated input text is deliberately *not* kept, in line with
CAMEL-25311's rule not to log submitted text.
** The buffer capacity is configurable and range-checked, also when it is set
after the console has started (see CAMEL-25346).
The recorder behind the ring buffer captures the same fields that the
OpenTelemetry {{gen_ai.evaluation.result}} event needs, and that the proposed
{{gen_ai.security.finding}} event (semantic-conventions-genai PR #427) needs
for guardrail-type experts. Emitting those events through the shared GenAI
observability module can reuse it (see the comments on CAMEL-25310 /
CAMEL-25311).
h3. 2. A camel-jbang screen backed by the console
A "Semantic" view, built like the other console-backed views:
{noformat}
Semantic ─ evaluations (live) ────────────────────────────────────────────
Time Route/Node Evaluation Expert Answer P Decision
10:42:07 ingest/choice2 injection security INJECTION 0.97 ▲ match
10:42:06 ingest/choice2 injection security benign 0.03 no match
10:42:05 triage/when1 department general billing 0.81 match
10:42:03 ingest/choice2 injection security benign 0.58 ? uncertain
─ Experts ──────────────────────────── ─ Declarations ────────────────────
security wolf-defender small-v2 injection BOOLEAN security t=0.5
BOOLEAN · no instructions department CHOICE general
general typesafe-ai jev-1.13.0 severity SCORE general
BOOLEAN/CHOICE/SCORE
{noformat}
* a detail popup per evaluation;
* a probability bar with a threshold marker;
* filtering by expert or evaluation name;
* positive security verdicts highlighted.
The mockup is a sketch. Expert names, providers and artifact identities are
illustrative.
h2. Out of scope
* Emitting the OpenTelemetry events (a separate follow-up once the recorder
exists; {{gen_ai.security.*}} is still an open proposal).
* Showing or storing evaluated input text.
_Filed by Claude Code on behalf of davsclaus_
--
This message was sent by Atlassian Jira
(v8.20.10#820010)