[
https://issues.apache.org/jira/browse/CAMEL-25382?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Luigi De Masi reassigned CAMEL-25382:
-------------------------------------
Assignee: Luigi De Masi
> camel-semantic: Introduce expert-owned evaluation contracts
> -----------------------------------------------------------
>
> Key: CAMEL-25382
> URL: https://issues.apache.org/jira/browse/CAMEL-25382
> Project: Camel
> Issue Type: Improvement
> Components: camel-ai, camel-core, dsl
> Reporter: Luigi De Masi
> Assignee: Luigi De Masi
> Priority: Major
>
> h2. Purpose
> Extend {{camel-semantic}} so that each expert defines the contract of the
> evaluations it supports: operations, accepted inputs, parameters and result
> semantics. Camel should provide the common declaration, validation and
> execution mechanisms, allowing routes to consume those results through
> existing EIPs.
> This is an initial design proposal for discussion. The examples illustrate
> proposed behaviour and syntax; they are not available APIs. The priorities
> are the exact contract definition and runtime/DSL integration. Tooling can
> follow separately.
> h2. Motivation
> Semantic evaluation should accommodate both instruction-driven experts, such
> as Jev-style decision services, and specialised models that assess security,
> content or response quality.
> These systems have different contracts:
> ||Example||Input and behaviour||Result||
> |Patronus Lion Warden|Evaluates text through several independent
> classification heads.|Binary detection, individual categories and multiple
> tags, with associated scores.|
> |Granite Guardian 4.1|Evaluates conversations against a criterion,
> potentially using supporting documents or tool definitions.|A yes/no
> judgement, optionally accompanied by reasoning.|
> |NVIDIA Nemotron Safety Guard v3|Evaluates user messages and optional
> assistant responses against a safety taxonomy.|Separate safety verdicts and
> applicable categories.|
> |Jev-style decision expert|Evaluates application instructions and criteria
> against supplied state.|A business decision, such as a department or urgency
> assessment.|
> References: [Patronus model
> card|https://huggingface.co/patronus-studio/lion-warden-ai-security-classifier],
> [Granite Guardian model
> card|https://huggingface.co/ibm-granite/granite-guardian-4.1-8b], and [NVIDIA
> model
> card|https://huggingface.co/nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3].
> The abstraction should support these differences without embedding a
> particular provider's taxonomy, request format or inference architecture into
> Camel.
> h2. Proposed ownership
> An expert remains an ordinary Camel registry bean implementing the semantic
> adapter SPI. Its bean name identifies the configured instance.
> The expert owns:
> * Supported operations and their meaning.
> * Accepted input shapes and requirements.
> * Evaluation parameters, their types and constraints.
> * Result types and their meaning.
> * Supported probability or confidence information.
> * Construction of provider requests and interpretation of responses.
> Camel owns:
> * Named evaluation declarations.
> * Expert resolution through the registry.
> * State selection.
> * Common contract validation.
> * Invocation and integration with expression languages and EIPs.
> * Publication of detailed semantic results.
> Endpoint addresses, credentials and deployment settings remain bean
> configuration. They are separate from the parameters of an evaluation.
> h2. Static contract
> Extend the existing {{@SemanticExpert}} annotation to describe operations and
> their contracts.
> ||Element||Description||
> |Identity|Operation name and purpose.|
> |Input|Accepted input shape and requirements.|
> |Parameters|Names, types, required values, constraints and documented
> omission behaviour.|
> |Result|Result type and meaning.|
> |Optional information|Supported probabilities, confidence and their
> interpretation.|
> The contract must be available without constructing an expert, loading a
> model or contacting a service.
> Configured instances must not redefine the static contract. Instance
> validation can reject an incompatible configuration, but cannot introduce
> operations or change result semantics.
> The same mechanism must describe instruction-driven experts. Requirements
> such as instructions, criteria or score levels belong to the relevant
> expert's contract.
> Future JSON descriptors should be generated from these annotations during the
> build and travel with the provider artifact. There should be no separately
> maintained contract definition. Descriptor generation and tooling integration
> are follow-up work and do not require a new Camel Catalog model or API in
> this proposal.
> h2. Evaluation declarations
> A semantic declaration defines named evaluations. A block can supply a common
> expert and state selector; individual evaluations can override them.
> The following example uses a nested {{parameters}} block as one syntax
> candidate. Expert-owned parameters are not required to be flat. The final
> parameter layout remains open.
> {noformat}
> - semantic:
> expert: security
> state: "${body}"
> evaluation:
> injection:
> operation: injection
> categories:
> operation: classify
> department:
> expert: decisions
> type: choice
> parameters:
> instructions: Which department should handle this message?
> criteria:
> billing: Payments, invoices and refunds
> technical: Errors, outages and technical questions
> sales: Product and purchasing questions
> {noformat}
> Here:
> * {{injection}}, {{categories}} and {{department}} are application-defined
> evaluation names.
> * {{security}} and {{decisions}} refer to registry beans.
> * Operation names belong to the expert contract.
> * The instruction-driven {{type: choice}} form is also governed by its
> expert's contract.
> * Declaring an evaluation does not execute it.
> The bean registrations could be supplied through {{application.properties}}:
> {noformat}
> camel.beans.security=#class:com.example.SecurityExpert
> camel.beans.decisions=#class:com.example.DecisionExpert
> {noformat}
> These class names and operation names are illustrative.
> h2. State and structured input
> Camel evaluates the state selector and passes its value unchanged to the
> expert.
> State may contain text or structured data accepted by the expert, such as a
> conversation, supporting documents or tool definitions. Camel should not
> impose a universal conversation format, flatten structured input into text or
> construct provider prompts.
> The expert validates the selected input and constructs the provider request
> without modifying the original state.
> h2. Results
> ||Type||Value exposed to the route||
> |Boolean|A Boolean with explicitly documented meaning.|
> |Choice|One category as a string.|
> |Score|A numeric value with a documented scale.|
> |Classification|A set of labels, potentially empty.|
> Classification is distinct from score. Labels are not represented as numeric
> scores.
> Result meaning must be explicit. For example, {{true}} could mean that a
> threat was detected or that a requirement was satisfied.
> Probability and confidence are optional. Their presence and interpretation
> depend on the operation contract. Camel must not manufacture confidence from
> a Boolean or categorical verdict.
> Where a provider returns several assessments, an adapter can expose them as
> separate operations. An operation need not correspond to a classifier head,
> HTTP endpoint or individual inference call. Any optimisation that obtains
> multiple results through one provider request remains an adapter
> responsibility.
> h2. Validation: shared responsibility
> The proposed approach divides validation between Camel and the expert:
> ||Stage||Camel||Expert||
> |Declaration/startup|Resolve expert and operation; validate recognised
> parameters, required values, types and declared constraints.|Validate
> operation-specific parameter combinations and compatibility with bean
> configuration.|
> |Before evaluation|Validate the selected state against the declared broad
> input shape.|Validate detailed content requirements, such as the presence of
> a response and supporting documents.|
> |After evaluation|Validate the returned type and common result
> constraints.|Parse and validate the provider response and translate it into
> the operation's documented meaning.|
> Validation should fail as early as the required information permits. Checks
> that need message data necessarily run during evaluation.
> Ordinary declaration validation should not perform inference or network
> calls. Diagnostics should identify the evaluation, expert and offending
> parameter without exposing credentials or message contents.
> Operational failures and malformed provider responses remain errors. They
> must not become negative or benign decisions.
> h2. Use from routes
> Introduce a Simple function:
> {noformat}
> ${semantic('evaluationName')}
> {noformat}
> The function invokes the named evaluation and returns its typed value. It
> preserves the message body and publishes the detailed result through the
> standard {{CamelSemanticResult}} exchange property.
> For example, detect injection and otherwise select a department using the
> instruction-driven evaluation above:
> {noformat}
> - route:
> from:
> uri: direct:incoming
> steps:
> - choice:
> when:
> - simple: "${semantic('injection')}"
> steps:
> - to: direct:security-review
> otherwise:
> steps:
> - switch:
> selector:
> simple: "${semantic('department')}"
> case:
> - value: billing
> uri: direct:billing
> - value: technical
> uri: direct:technical
> - value: sales
> uri: direct:sales
> otherwise:
> uri: direct:manual-review
> {noformat}
> Classification results can be used in membership predicates:
> {noformat}
> - choice:
> when:
> - simple: "${semantic('categories')} contains 'privacy'"
> steps:
> - to: direct:privacy-review
> otherwise:
> steps:
> - to: direct:continue
> {noformat}
> The category is illustrative and must be meaningful for the selected expert.
> Each function evaluation invokes the expert. There is no implicit caching
> across occurrences. Applications needing reuse can explicitly retain the
> returned value in a Camel variable.
> Switch continues to dispatch scalar results to route-defined destinations.
> Collections are consumed through predicates such as membership checks.
> Evaluation failures follow normal Camel error handling.
> h2. Non-Jev example: specialised content moderation
> An expert backed by NVIDIA Nemotron Safety Guard can assess user messages and
> assistant responses. The model returns separate safety verdicts and
> applicable categories; see the [model
> documentation|https://huggingface.co/nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3#quick-start].
> The following adapter class, configuration options and operation names are
> illustrative proposals.
> Register the expert through {{application.properties}}:
> {noformat}
> camel.beans.moderation=#class:com.example.NemotronSafetyExpert
> camel.beans.moderation.baseUrl={{env:MODERATION_URL}}
> {noformat}
> The proposed expert accepts structured state containing a user message and an
> optional assistant response. For example, the message body could contain a
> map equivalent to:
> {noformat}
> {
> "prompt": "Give me my coworker's private home address.",
> "response": "I cannot help disclose someone's private information."
> }
> {noformat}
> Camel passes this map unchanged. The adapter constructs the model-specific
> request.
> The expert contract exposes:
> ||Operation||Result type||Meaning||
> |{{user_safety}}|Choice|Safety verdict for the user message.|
> |{{response_safety}}|Choice|Safety verdict for the assistant response,
> considering the supplied context.|
> |{{categories}}|Classification|Applicable categories across the assessed
> content.|
> Applications select these operations through named evaluations:
> {noformat}
> - semantic:
> expert: moderation
> state: "${body}"
> evaluation:
> userSafety:
> operation: user_safety
> responseSafety:
> operation: response_safety
> safetyCategories:
> operation: categories
> {noformat}
> The expert supplies the evaluation semantics and result types. No
> application-written instructions or criteria are required.
> A route can dispatch an assessed assistant response using Switch:
> {noformat}
> - route:
> from:
> uri: direct:check-response
> steps:
> - switch:
> selector:
> simple: "${semantic('responseSafety')}"
> case:
> - value: safe
> uri: direct:deliver-response
> - value: unsafe
> uri: direct:review-response
> otherwise:
> uri: direct:manual-review
> {noformat}
> The response assessment requires an assistant response in the selected state.
> The expert validates that requirement before invoking the service.
> Provider failures and malformed verdicts follow Camel error handling.
> Detailed results remain available through {{CamelSemanticResult}}.
> This example demonstrates structured input, expert-owned operations and a
> categorical result used directly in routing. The Jev-style example
> demonstrates application-supplied decision criteria through the same semantic
> infrastructure.
> h2. Runtime and DSL integration
> Implementation should cover:
> * Operation metadata in {{@SemanticExpert}} and its immutable runtime
> representation.
> * A common evaluation declaration supporting expert-owned parameters,
> including nested values.
> * Adapter validation and execution against the resolved contract.
> * Classification results and optional associated confidence information.
> * Equivalent declaration capabilities in YAML, XML and Java.
> * The Simple function and continued semantic-language invocation, including
> {{ref:name}}.
> * Consistent result publication, including removal of stale results when an
> invocation fails.
> YAML, XML and Java should converge on the same declaration model and
> validation path. Their surface syntax does not need to be identical.
> h2. Open questions
> # *Parameter representation:* should expert parameters use a dedicated
> container, direct fields, or another representation? How should nested maps
> and lists be expressed consistently across DSLs?
> # *Decision policy:* which layer applies explicitly requested thresholds -
> the service, adapter or Camel?
> # *Omitted policy:* how should service or adapter defaults be preserved when
> the route supplies no threshold?
> # *Missing confidence:* how should a requested confidence-dependent policy
> behave when the operation or a particular response cannot supply the required
> information?
> # *Uncertain outcomes:* how should they be represented and handled?
> # *Classification filtering:* how should a per-label confidence threshold be
> expressed and applied?
> The proposal must make policy ownership explicit so that thresholds are not
> silently applied more than once.
> h2. Scope and compatibility
> The previous implementation has not been released, so the semantic API, SPI
> and declaration syntax can be revised without a backward-compatibility layer.
> The initial scope is the exact contract definition and runtime/DSL
> integration.
> Camel Catalog model and API changes are excluded from this proposal. IDE
> completion, Kaoto integration, offline expert-specific validation and
> generated contract descriptors can follow separately. Coordinate
> declaration/tooling work with CAMEL-25259 and CAMEL-25257; this proposal does
> not replace those efforts or introduce their catalog changes.
> Model weights and provider runtimes do not need to be bundled with Camel.
> Provider adapters connect to the relevant services or implementations.
> Implementing every example provider is not a prerequisite for agreeing on and
> validating the common contract.
> h2. Acceptance criteria
> # Contracts can be inspected without constructing or contacting an expert.
> # A configured instance cannot redefine its static contract.
> # Instruction-driven and specialised experts use the same evaluation
> infrastructure.
> # Invalid declarations fail before processing messages whenever the necessary
> information is available.
> # Structured state reaches the expert unchanged.
> # Boolean, choice, score and classification results are validated and usable
> in routes.
> # Provider errors and malformed results follow Camel error handling.
> # YAML, XML and Java have equivalent declaration semantics and shared
> validation.
> # Tests cover representative text and structured-input experts without
> requiring live model services.
> # Documentation explains expert configuration, result meaning, invocation
> behaviour and the agreed decision policy, with both instruction-driven and
> specialised examples.
> h2. Related work
> * CAMEL-25310 provides the semantic expert foundation.
> * CAMEL-25311 is a potential consumer of the extended contract. Its provider
> implementation remains a separate issue.
> * CAMEL-25259 discusses DSL-extension models and tooling metadata.
> * CAMEL-25257 discusses dependency detection and DSL conversion for feature
> declarations.
> * CAMEL-25357 discusses a semantic dev console and can consume the resulting
> contracts separately.
> _AI-generated proposal by Codex on behalf of
> [luigidemasi|https://github.com/luigidemasi]._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)