Luigi De Masi created CAMEL-25382:
-------------------------------------
Summary: camel-semantic: Introduce expert-owned evaluation
contracts
Key: CAMEL-25382
URL: https://issues.apache.org/jira/browse/CAMEL-25382
Project: Camel
Issue Type: Improvement
Components: dsl, camel-core, camel-ai
Reporter: Luigi De Masi
h2. Purpose
Extend {{camel-semantic}} so that each expert defines the contract of the
evaluations it supports: operations, accepted inputs, parameters and result
semantics. Camel should provide the common declaration, validation and
execution mechanisms, allowing routes to consume those results through existing
EIPs.
This is an initial design proposal for discussion. The examples illustrate
proposed behaviour and syntax; they are not available APIs. The priorities are
the exact contract definition and runtime/DSL integration. Tooling can follow
separately.
h2. Motivation
Semantic evaluation should accommodate both instruction-driven experts, such as
Jev-style decision services, and specialised models that assess security,
content or response quality.
These systems have different contracts:
||Example||Input and behaviour||Result||
|Patronus Lion Warden|Evaluates text through several independent classification
heads.|Binary detection, individual categories and multiple tags, with
associated scores.|
|Granite Guardian 4.1|Evaluates conversations against a criterion, potentially
using supporting documents or tool definitions.|A yes/no judgement, optionally
accompanied by reasoning.|
|NVIDIA Nemotron Safety Guard v3|Evaluates user messages and optional assistant
responses against a safety taxonomy.|Separate safety verdicts and applicable
categories.|
|Jev-style decision expert|Evaluates application instructions and criteria
against supplied state.|A business decision, such as a department or urgency
assessment.|
References: [Patronus model
card|https://huggingface.co/patronus-studio/lion-warden-ai-security-classifier],
[Granite Guardian model
card|https://huggingface.co/ibm-granite/granite-guardian-4.1-8b], and [NVIDIA
model card|https://huggingface.co/nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3].
The abstraction should support these differences without embedding a particular
provider's taxonomy, request format or inference architecture into Camel.
h2. Proposed ownership
An expert remains an ordinary Camel registry bean implementing the semantic
adapter SPI. Its bean name identifies the configured instance.
The expert owns:
* Supported operations and their meaning.
* Accepted input shapes and requirements.
* Evaluation parameters, their types and constraints.
* Result types and their meaning.
* Supported probability or confidence information.
* Construction of provider requests and interpretation of responses.
Camel owns:
* Named evaluation declarations.
* Expert resolution through the registry.
* State selection.
* Common contract validation.
* Invocation and integration with expression languages and EIPs.
* Publication of detailed semantic results.
Endpoint addresses, credentials and deployment settings remain bean
configuration. They are separate from the parameters of an evaluation.
h2. Static contract
Extend the existing {{@SemanticExpert}} annotation to describe operations and
their contracts.
||Element||Description||
|Identity|Operation name and purpose.|
|Input|Accepted input shape and requirements.|
|Parameters|Names, types, required values, constraints and documented omission
behaviour.|
|Result|Result type and meaning.|
|Optional information|Supported probabilities, confidence and their
interpretation.|
The contract must be available without constructing an expert, loading a model
or contacting a service.
Configured instances must not redefine the static contract. Instance validation
can reject an incompatible configuration, but cannot introduce operations or
change result semantics.
The same mechanism must describe instruction-driven experts. Requirements such
as instructions, criteria or score levels belong to the relevant expert's
contract.
Future JSON descriptors should be generated from these annotations during the
build and travel with the provider artifact. There should be no separately
maintained contract definition. Descriptor generation and tooling integration
are follow-up work and do not require a new Camel Catalog model or API in this
proposal.
h2. Evaluation declarations
A semantic declaration defines named evaluations. A block can supply a common
expert and state selector; individual evaluations can override them.
The following example uses a nested {{parameters}} block as one syntax
candidate. Expert-owned parameters are not required to be flat. The final
parameter layout remains open.
{noformat}
- semantic:
expert: security
state: "${body}"
evaluation:
injection:
operation: injection
categories:
operation: classify
department:
expert: decisions
type: choice
parameters:
instructions: Which department should handle this message?
criteria:
billing: Payments, invoices and refunds
technical: Errors, outages and technical questions
sales: Product and purchasing questions
{noformat}
Here:
* {{injection}}, {{categories}} and {{department}} are application-defined
evaluation names.
* {{security}} and {{decisions}} refer to registry beans.
* Operation names belong to the expert contract.
* The instruction-driven {{type: choice}} form is also governed by its expert's
contract.
* Declaring an evaluation does not execute it.
The bean registrations could be supplied through {{application.properties}}:
{noformat}
camel.beans.security=#class:com.example.SecurityExpert
camel.beans.decisions=#class:com.example.DecisionExpert
{noformat}
These class names and operation names are illustrative.
h2. State and structured input
Camel evaluates the state selector and passes its value unchanged to the expert.
State may contain text or structured data accepted by the expert, such as a
conversation, supporting documents or tool definitions. Camel should not impose
a universal conversation format, flatten structured input into text or
construct provider prompts.
The expert validates the selected input and constructs the provider request
without modifying the original state.
h2. Results
||Type||Value exposed to the route||
|Boolean|A Boolean with explicitly documented meaning.|
|Choice|One category as a string.|
|Score|A numeric value with a documented scale.|
|Classification|A set of labels, potentially empty.|
Classification is distinct from score. Labels are not represented as numeric
scores.
Result meaning must be explicit. For example, {{true}} could mean that a threat
was detected or that a requirement was satisfied.
Probability and confidence are optional. Their presence and interpretation
depend on the operation contract. Camel must not manufacture confidence from a
Boolean or categorical verdict.
Where a provider returns several assessments, an adapter can expose them as
separate operations. An operation need not correspond to a classifier head,
HTTP endpoint or individual inference call. Any optimisation that obtains
multiple results through one provider request remains an adapter responsibility.
h2. Validation: shared responsibility
The proposed approach divides validation between Camel and the expert:
||Stage||Camel||Expert||
|Declaration/startup|Resolve expert and operation; validate recognised
parameters, required values, types and declared constraints.|Validate
operation-specific parameter combinations and compatibility with bean
configuration.|
|Before evaluation|Validate the selected state against the declared broad input
shape.|Validate detailed content requirements, such as the presence of a
response and supporting documents.|
|After evaluation|Validate the returned type and common result
constraints.|Parse and validate the provider response and translate it into the
operation's documented meaning.|
Validation should fail as early as the required information permits. Checks
that need message data necessarily run during evaluation.
Ordinary declaration validation should not perform inference or network calls.
Diagnostics should identify the evaluation, expert and offending parameter
without exposing credentials or message contents.
Operational failures and malformed provider responses remain errors. They must
not become negative or benign decisions.
h2. Use from routes
Introduce a Simple function:
{noformat}
${semantic('evaluationName')}
{noformat}
The function invokes the named evaluation and returns its typed value. It
preserves the message body and publishes the detailed result through the
standard {{CamelSemanticResult}} exchange property.
For example, detect injection and otherwise select a department using the
instruction-driven evaluation above:
{noformat}
- route:
from:
uri: direct:incoming
steps:
- choice:
when:
- simple: "${semantic('injection')}"
steps:
- to: direct:security-review
otherwise:
steps:
- switch:
selector:
simple: "${semantic('department')}"
case:
- value: billing
uri: direct:billing
- value: technical
uri: direct:technical
- value: sales
uri: direct:sales
otherwise:
uri: direct:manual-review
{noformat}
Classification results can be used in membership predicates:
{noformat}
- choice:
when:
- simple: "${semantic('categories')} contains 'privacy'"
steps:
- to: direct:privacy-review
otherwise:
steps:
- to: direct:continue
{noformat}
The category is illustrative and must be meaningful for the selected expert.
Each function evaluation invokes the expert. There is no implicit caching
across occurrences. Applications needing reuse can explicitly retain the
returned value in a Camel variable.
Switch continues to dispatch scalar results to route-defined destinations.
Collections are consumed through predicates such as membership checks.
Evaluation failures follow normal Camel error handling.
h2. Non-Jev example: specialised content moderation
An expert backed by NVIDIA Nemotron Safety Guard can assess user messages and
assistant responses. The model returns separate safety verdicts and applicable
categories; see the [model
documentation|https://huggingface.co/nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3#quick-start].
The following adapter class, configuration options and operation names are
illustrative proposals.
Register the expert through {{application.properties}}:
{noformat}
camel.beans.moderation=#class:com.example.NemotronSafetyExpert
camel.beans.moderation.baseUrl={{env:MODERATION_URL}}
{noformat}
The proposed expert accepts structured state containing a user message and an
optional assistant response. For example, the message body could contain a map
equivalent to:
{noformat}
{
"prompt": "Give me my coworker's private home address.",
"response": "I cannot help disclose someone's private information."
}
{noformat}
Camel passes this map unchanged. The adapter constructs the model-specific
request.
The expert contract exposes:
||Operation||Result type||Meaning||
|{{user_safety}}|Choice|Safety verdict for the user message.|
|{{response_safety}}|Choice|Safety verdict for the assistant response,
considering the supplied context.|
|{{categories}}|Classification|Applicable categories across the assessed
content.|
Applications select these operations through named evaluations:
{noformat}
- semantic:
expert: moderation
state: "${body}"
evaluation:
userSafety:
operation: user_safety
responseSafety:
operation: response_safety
safetyCategories:
operation: categories
{noformat}
The expert supplies the evaluation semantics and result types. No
application-written instructions or criteria are required.
A route can dispatch an assessed assistant response using Switch:
{noformat}
- route:
from:
uri: direct:check-response
steps:
- switch:
selector:
simple: "${semantic('responseSafety')}"
case:
- value: safe
uri: direct:deliver-response
- value: unsafe
uri: direct:review-response
otherwise:
uri: direct:manual-review
{noformat}
The response assessment requires an assistant response in the selected state.
The expert validates that requirement before invoking the service.
Provider failures and malformed verdicts follow Camel error handling. Detailed
results remain available through {{CamelSemanticResult}}.
This example demonstrates structured input, expert-owned operations and a
categorical result used directly in routing. The Jev-style example demonstrates
application-supplied decision criteria through the same semantic infrastructure.
h2. Runtime and DSL integration
Implementation should cover:
* Operation metadata in {{@SemanticExpert}} and its immutable runtime
representation.
* A common evaluation declaration supporting expert-owned parameters, including
nested values.
* Adapter validation and execution against the resolved contract.
* Classification results and optional associated confidence information.
* Equivalent declaration capabilities in YAML, XML and Java.
* The Simple function and continued semantic-language invocation, including
{{ref:name}}.
* Consistent result publication, including removal of stale results when an
invocation fails.
YAML, XML and Java should converge on the same declaration model and validation
path. Their surface syntax does not need to be identical.
h2. Open questions
# *Parameter representation:* should expert parameters use a dedicated
container, direct fields, or another representation? How should nested maps and
lists be expressed consistently across DSLs?
# *Decision policy:* which layer applies explicitly requested thresholds - the
service, adapter or Camel?
# *Omitted policy:* how should service or adapter defaults be preserved when
the route supplies no threshold?
# *Missing confidence:* how should a requested confidence-dependent policy
behave when the operation or a particular response cannot supply the required
information?
# *Uncertain outcomes:* how should they be represented and handled?
# *Classification filtering:* how should a per-label confidence threshold be
expressed and applied?
The proposal must make policy ownership explicit so that thresholds are not
silently applied more than once.
h2. Scope and compatibility
The previous implementation has not been released, so the semantic API, SPI and
declaration syntax can be revised without a backward-compatibility layer.
The initial scope is the exact contract definition and runtime/DSL integration.
Camel Catalog model and API changes are excluded from this proposal. IDE
completion, Kaoto integration, offline expert-specific validation and generated
contract descriptors can follow separately. Coordinate declaration/tooling work
with CAMEL-25259 and CAMEL-25257; this proposal does not replace those efforts
or introduce their catalog changes.
Model weights and provider runtimes do not need to be bundled with Camel.
Provider adapters connect to the relevant services or implementations.
Implementing every example provider is not a prerequisite for agreeing on and
validating the common contract.
h2. Acceptance criteria
# Contracts can be inspected without constructing or contacting an expert.
# A configured instance cannot redefine its static contract.
# Instruction-driven and specialised experts use the same evaluation
infrastructure.
# Invalid declarations fail before processing messages whenever the necessary
information is available.
# Structured state reaches the expert unchanged.
# Boolean, choice, score and classification results are validated and usable in
routes.
# Provider errors and malformed results follow Camel error handling.
# YAML, XML and Java have equivalent declaration semantics and shared
validation.
# Tests cover representative text and structured-input experts without
requiring live model services.
# Documentation explains expert configuration, result meaning, invocation
behaviour and the agreed decision policy, with both instruction-driven and
specialised examples.
h2. Related work
* CAMEL-25310 provides the semantic expert foundation.
* CAMEL-25311 is a potential consumer of the extended contract. Its provider
implementation remains a separate issue.
* CAMEL-25259 discusses DSL-extension models and tooling metadata.
* CAMEL-25257 discusses dependency detection and DSL conversion for feature
declarations.
* CAMEL-25357 discusses a semantic dev console and can consume the resulting
contracts separately.
_AI-generated proposal by Codex on behalf of
[luigidemasi|https://github.com/luigidemasi]._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)