Claus Ibsen created CAMEL-25357:
-----------------------------------

             Summary: camel-semantic - dev console for semantic evaluations and 
experts
                 Key: CAMEL-25357
                 URL: https://issues.apache.org/jira/browse/CAMEL-25357
             Project: Camel
          Issue Type: New Feature
          Components: camel-jbang, camel-ai
            Reporter: Claus Ibsen


cc [~ldemasi]

h2. Motivation

With experts and capabilities from CAMEL-25310 (and fixed-purpose experts such 
as the prompt-injection expert in CAMEL-25311), a running Camel application 
makes semantic decisions in its routes. Today nothing shows which evaluations 
are declared, which expert each one uses, or what the recent evaluations 
returned. camel-semantic has no dev console and emits no observability events.

This issue depends on CAMEL-25310 (PR #27373), since experts and 
{{SemanticCapabilities}} come from there.

h2. Proposal

h3. 1. A {{semantic}} dev console in camel-semantic

* *Declarations*: name, result type (BOOLEAN / CHOICE / SCORE), selected 
expert, threshold and uncertainty policy, whether instructions are set.
* *Experts*: configured instances, provider and pinned artifact/model identity, 
and their {{SemanticCapabilities}} (supported result types, instructions 
required/optional/unsupported, fixed meaning such as "positive = prompt 
injection detected", meaning of probability/confidence).
* *Recent evaluations*: a bounded ring buffer. Each entry has the time, 
route/node, evaluation name, expert, answer/label, probability, confidence, 
threshold decision (match / no match / uncertain), latency and error.
** The evaluated input text is deliberately *not* kept, in line with 
CAMEL-25311's rule not to log submitted text.
** The buffer capacity is configurable and range-checked, also when it is set 
after the console has started (see CAMEL-25346).

The recorder behind the ring buffer captures the same fields that the 
OpenTelemetry {{gen_ai.evaluation.result}} event needs, and that the proposed 
{{gen_ai.security.finding}} event (semantic-conventions-genai PR #427) needs 
for guardrail-type experts. Emitting those events through the shared GenAI 
observability module can reuse it (see the comments on CAMEL-25310 / 
CAMEL-25311).

h3. 2. A camel-jbang screen backed by the console

A "Semantic" view, built like the other console-backed views:

{noformat}
 Semantic ─ evaluations (live) ────────────────────────────────────────────
 Time      Route/Node        Evaluation  Expert    Answer     P     Decision
 10:42:07  ingest/choice2    injection   security  INJECTION  0.97  ▲ match
 10:42:06  ingest/choice2    injection   security  benign     0.03    no match
 10:42:05  triage/when1      department  general   billing    0.81    match
 10:42:03  ingest/choice2    injection   security  benign     0.58  ? uncertain
 ─ Experts ──────────────────────────── ─ Declarations ────────────────────
 security  wolf-defender small-v2       injection   BOOLEAN  security  t=0.5
           BOOLEAN · no instructions    department  CHOICE   general
 general   typesafe-ai jev-1.13.0       severity    SCORE    general
           BOOLEAN/CHOICE/SCORE
{noformat}

* a detail popup per evaluation;
* a probability bar with a threshold marker;
* filtering by expert or evaluation name;
* positive security verdicts highlighted.

The mockup is a sketch. Expert names, providers and artifact identities are 
illustrative.

h2. Out of scope

* Emitting the OpenTelemetry events (a separate follow-up once the recorder 
exists; {{gen_ai.security.*}} is still an open proposal).
* Showing or storing evaluated input text.

_Filed by Claude Code on behalf of davsclaus_




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to