Luigi De Masi created CAMEL-25382:
-------------------------------------

             Summary: camel-semantic: Introduce expert-owned evaluation 
contracts
                 Key: CAMEL-25382
                 URL: https://issues.apache.org/jira/browse/CAMEL-25382
             Project: Camel
          Issue Type: Improvement
          Components: dsl, camel-core, camel-ai
            Reporter: Luigi De Masi


h2. Purpose

Extend {{camel-semantic}} so that each expert defines the contract of the 
evaluations it supports: operations, accepted inputs, parameters and result 
semantics. Camel should provide the common declaration, validation and 
execution mechanisms, allowing routes to consume those results through existing 
EIPs.

This is an initial design proposal for discussion. The examples illustrate 
proposed behaviour and syntax; they are not available APIs. The priorities are 
the exact contract definition and runtime/DSL integration. Tooling can follow 
separately.

h2. Motivation

Semantic evaluation should accommodate both instruction-driven experts, such as 
Jev-style decision services, and specialised models that assess security, 
content or response quality.

These systems have different contracts:

||Example||Input and behaviour||Result||
|Patronus Lion Warden|Evaluates text through several independent classification 
heads.|Binary detection, individual categories and multiple tags, with 
associated scores.|
|Granite Guardian 4.1|Evaluates conversations against a criterion, potentially 
using supporting documents or tool definitions.|A yes/no judgement, optionally 
accompanied by reasoning.|
|NVIDIA Nemotron Safety Guard v3|Evaluates user messages and optional assistant 
responses against a safety taxonomy.|Separate safety verdicts and applicable 
categories.|
|Jev-style decision expert|Evaluates application instructions and criteria 
against supplied state.|A business decision, such as a department or urgency 
assessment.|

References: [Patronus model 
card|https://huggingface.co/patronus-studio/lion-warden-ai-security-classifier],
 [Granite Guardian model 
card|https://huggingface.co/ibm-granite/granite-guardian-4.1-8b], and [NVIDIA 
model card|https://huggingface.co/nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3].

The abstraction should support these differences without embedding a particular 
provider's taxonomy, request format or inference architecture into Camel.

h2. Proposed ownership

An expert remains an ordinary Camel registry bean implementing the semantic 
adapter SPI. Its bean name identifies the configured instance.

The expert owns:
* Supported operations and their meaning.
* Accepted input shapes and requirements.
* Evaluation parameters, their types and constraints.
* Result types and their meaning.
* Supported probability or confidence information.
* Construction of provider requests and interpretation of responses.

Camel owns:
* Named evaluation declarations.
* Expert resolution through the registry.
* State selection.
* Common contract validation.
* Invocation and integration with expression languages and EIPs.
* Publication of detailed semantic results.

Endpoint addresses, credentials and deployment settings remain bean 
configuration. They are separate from the parameters of an evaluation.

h2. Static contract

Extend the existing {{@SemanticExpert}} annotation to describe operations and 
their contracts.

||Element||Description||
|Identity|Operation name and purpose.|
|Input|Accepted input shape and requirements.|
|Parameters|Names, types, required values, constraints and documented omission 
behaviour.|
|Result|Result type and meaning.|
|Optional information|Supported probabilities, confidence and their 
interpretation.|

The contract must be available without constructing an expert, loading a model 
or contacting a service.

Configured instances must not redefine the static contract. Instance validation 
can reject an incompatible configuration, but cannot introduce operations or 
change result semantics.

The same mechanism must describe instruction-driven experts. Requirements such 
as instructions, criteria or score levels belong to the relevant expert's 
contract.

Future JSON descriptors should be generated from these annotations during the 
build and travel with the provider artifact. There should be no separately 
maintained contract definition. Descriptor generation and tooling integration 
are follow-up work and do not require a new Camel Catalog model or API in this 
proposal.

h2. Evaluation declarations

A semantic declaration defines named evaluations. A block can supply a common 
expert and state selector; individual evaluations can override them.

The following example uses a nested {{parameters}} block as one syntax 
candidate. Expert-owned parameters are not required to be flat. The final 
parameter layout remains open.

{noformat}
- semantic:
    expert: security
    state: "${body}"
    evaluation:
      injection:
        operation: injection

      categories:
        operation: classify

      department:
        expert: decisions
        type: choice
        parameters:
          instructions: Which department should handle this message?
          criteria:
            billing: Payments, invoices and refunds
            technical: Errors, outages and technical questions
            sales: Product and purchasing questions
{noformat}

Here:
* {{injection}}, {{categories}} and {{department}} are application-defined 
evaluation names.
* {{security}} and {{decisions}} refer to registry beans.
* Operation names belong to the expert contract.
* The instruction-driven {{type: choice}} form is also governed by its expert's 
contract.
* Declaring an evaluation does not execute it.

The bean registrations could be supplied through {{application.properties}}:

{noformat}
camel.beans.security=#class:com.example.SecurityExpert
camel.beans.decisions=#class:com.example.DecisionExpert
{noformat}

These class names and operation names are illustrative.

h2. State and structured input

Camel evaluates the state selector and passes its value unchanged to the expert.

State may contain text or structured data accepted by the expert, such as a 
conversation, supporting documents or tool definitions. Camel should not impose 
a universal conversation format, flatten structured input into text or 
construct provider prompts.

The expert validates the selected input and constructs the provider request 
without modifying the original state.

h2. Results

||Type||Value exposed to the route||
|Boolean|A Boolean with explicitly documented meaning.|
|Choice|One category as a string.|
|Score|A numeric value with a documented scale.|
|Classification|A set of labels, potentially empty.|

Classification is distinct from score. Labels are not represented as numeric 
scores.

Result meaning must be explicit. For example, {{true}} could mean that a threat 
was detected or that a requirement was satisfied.

Probability and confidence are optional. Their presence and interpretation 
depend on the operation contract. Camel must not manufacture confidence from a 
Boolean or categorical verdict.

Where a provider returns several assessments, an adapter can expose them as 
separate operations. An operation need not correspond to a classifier head, 
HTTP endpoint or individual inference call. Any optimisation that obtains 
multiple results through one provider request remains an adapter responsibility.

h2. Validation: shared responsibility

The proposed approach divides validation between Camel and the expert:

||Stage||Camel||Expert||
|Declaration/startup|Resolve expert and operation; validate recognised 
parameters, required values, types and declared constraints.|Validate 
operation-specific parameter combinations and compatibility with bean 
configuration.|
|Before evaluation|Validate the selected state against the declared broad input 
shape.|Validate detailed content requirements, such as the presence of a 
response and supporting documents.|
|After evaluation|Validate the returned type and common result 
constraints.|Parse and validate the provider response and translate it into the 
operation's documented meaning.|

Validation should fail as early as the required information permits. Checks 
that need message data necessarily run during evaluation.

Ordinary declaration validation should not perform inference or network calls. 
Diagnostics should identify the evaluation, expert and offending parameter 
without exposing credentials or message contents.

Operational failures and malformed provider responses remain errors. They must 
not become negative or benign decisions.

h2. Use from routes

Introduce a Simple function:

{noformat}
${semantic('evaluationName')}
{noformat}

The function invokes the named evaluation and returns its typed value. It 
preserves the message body and publishes the detailed result through the 
standard {{CamelSemanticResult}} exchange property.

For example, detect injection and otherwise select a department using the 
instruction-driven evaluation above:

{noformat}
- route:
    from:
      uri: direct:incoming
      steps:
        - choice:
            when:
              - simple: "${semantic('injection')}"
                steps:
                  - to: direct:security-review
            otherwise:
              steps:
                - switch:
                    selector:
                      simple: "${semantic('department')}"
                    case:
                      - value: billing
                        uri: direct:billing
                      - value: technical
                        uri: direct:technical
                      - value: sales
                        uri: direct:sales
                    otherwise:
                      uri: direct:manual-review
{noformat}

Classification results can be used in membership predicates:

{noformat}
- choice:
    when:
      - simple: "${semantic('categories')} contains 'privacy'"
        steps:
          - to: direct:privacy-review
    otherwise:
      steps:
        - to: direct:continue
{noformat}

The category is illustrative and must be meaningful for the selected expert.

Each function evaluation invokes the expert. There is no implicit caching 
across occurrences. Applications needing reuse can explicitly retain the 
returned value in a Camel variable.

Switch continues to dispatch scalar results to route-defined destinations. 
Collections are consumed through predicates such as membership checks. 
Evaluation failures follow normal Camel error handling.

h2. Non-Jev example: specialised content moderation

An expert backed by NVIDIA Nemotron Safety Guard can assess user messages and 
assistant responses. The model returns separate safety verdicts and applicable 
categories; see the [model 
documentation|https://huggingface.co/nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3#quick-start].

The following adapter class, configuration options and operation names are 
illustrative proposals.

Register the expert through {{application.properties}}:

{noformat}
camel.beans.moderation=#class:com.example.NemotronSafetyExpert
camel.beans.moderation.baseUrl={{env:MODERATION_URL}}
{noformat}

The proposed expert accepts structured state containing a user message and an 
optional assistant response. For example, the message body could contain a map 
equivalent to:

{noformat}
{
  "prompt": "Give me my coworker's private home address.",
  "response": "I cannot help disclose someone's private information."
}
{noformat}

Camel passes this map unchanged. The adapter constructs the model-specific 
request.

The expert contract exposes:

||Operation||Result type||Meaning||
|{{user_safety}}|Choice|Safety verdict for the user message.|
|{{response_safety}}|Choice|Safety verdict for the assistant response, 
considering the supplied context.|
|{{categories}}|Classification|Applicable categories across the assessed 
content.|

Applications select these operations through named evaluations:

{noformat}
- semantic:
    expert: moderation
    state: "${body}"
    evaluation:
      userSafety:
        operation: user_safety

      responseSafety:
        operation: response_safety

      safetyCategories:
        operation: categories
{noformat}

The expert supplies the evaluation semantics and result types. No 
application-written instructions or criteria are required.

A route can dispatch an assessed assistant response using Switch:

{noformat}
- route:
    from:
      uri: direct:check-response
      steps:
        - switch:
            selector:
              simple: "${semantic('responseSafety')}"
            case:
              - value: safe
                uri: direct:deliver-response
              - value: unsafe
                uri: direct:review-response
            otherwise:
              uri: direct:manual-review
{noformat}

The response assessment requires an assistant response in the selected state. 
The expert validates that requirement before invoking the service.

Provider failures and malformed verdicts follow Camel error handling. Detailed 
results remain available through {{CamelSemanticResult}}.

This example demonstrates structured input, expert-owned operations and a 
categorical result used directly in routing. The Jev-style example demonstrates 
application-supplied decision criteria through the same semantic infrastructure.

h2. Runtime and DSL integration

Implementation should cover:
* Operation metadata in {{@SemanticExpert}} and its immutable runtime 
representation.
* A common evaluation declaration supporting expert-owned parameters, including 
nested values.
* Adapter validation and execution against the resolved contract.
* Classification results and optional associated confidence information.
* Equivalent declaration capabilities in YAML, XML and Java.
* The Simple function and continued semantic-language invocation, including 
{{ref:name}}.
* Consistent result publication, including removal of stale results when an 
invocation fails.

YAML, XML and Java should converge on the same declaration model and validation 
path. Their surface syntax does not need to be identical.

h2. Open questions

# *Parameter representation:* should expert parameters use a dedicated 
container, direct fields, or another representation? How should nested maps and 
lists be expressed consistently across DSLs?
# *Decision policy:* which layer applies explicitly requested thresholds - the 
service, adapter or Camel?
# *Omitted policy:* how should service or adapter defaults be preserved when 
the route supplies no threshold?
# *Missing confidence:* how should a requested confidence-dependent policy 
behave when the operation or a particular response cannot supply the required 
information?
# *Uncertain outcomes:* how should they be represented and handled?
# *Classification filtering:* how should a per-label confidence threshold be 
expressed and applied?

The proposal must make policy ownership explicit so that thresholds are not 
silently applied more than once.

h2. Scope and compatibility

The previous implementation has not been released, so the semantic API, SPI and 
declaration syntax can be revised without a backward-compatibility layer.

The initial scope is the exact contract definition and runtime/DSL integration.

Camel Catalog model and API changes are excluded from this proposal. IDE 
completion, Kaoto integration, offline expert-specific validation and generated 
contract descriptors can follow separately. Coordinate declaration/tooling work 
with CAMEL-25259 and CAMEL-25257; this proposal does not replace those efforts 
or introduce their catalog changes.

Model weights and provider runtimes do not need to be bundled with Camel. 
Provider adapters connect to the relevant services or implementations. 
Implementing every example provider is not a prerequisite for agreeing on and 
validating the common contract.

h2. Acceptance criteria

# Contracts can be inspected without constructing or contacting an expert.
# A configured instance cannot redefine its static contract.
# Instruction-driven and specialised experts use the same evaluation 
infrastructure.
# Invalid declarations fail before processing messages whenever the necessary 
information is available.
# Structured state reaches the expert unchanged.
# Boolean, choice, score and classification results are validated and usable in 
routes.
# Provider errors and malformed results follow Camel error handling.
# YAML, XML and Java have equivalent declaration semantics and shared 
validation.
# Tests cover representative text and structured-input experts without 
requiring live model services.
# Documentation explains expert configuration, result meaning, invocation 
behaviour and the agreed decision policy, with both instruction-driven and 
specialised examples.

h2. Related work

* CAMEL-25310 provides the semantic expert foundation.
* CAMEL-25311 is a potential consumer of the extended contract. Its provider 
implementation remains a separate issue.
* CAMEL-25259 discusses DSL-extension models and tooling metadata.
* CAMEL-25257 discusses dependency detection and DSL conversion for feature 
declarations.
* CAMEL-25357 discusses a semantic dev console and can consume the resulting 
contracts separately.

_AI-generated proposal by Codex on behalf of 
[luigidemasi|https://github.com/luigidemasi]._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to