[ 
https://issues.apache.org/jira/browse/CAMEL-25311?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18123470#comment-18123470
 ] 

Claus Ibsen commented on CAMEL-25311:
-------------------------------------

Related to observability for a guardrail-type expert like this one, and not a 
blocker for the initial implementation.

OpenTelemetry is working on guardrail-specific conventions: 
semantic-conventions-genai PR #427 "run guardrail span and security finding" 
(still open, so names may change): 
https://github.com/open-telemetry/semantic-conventions-genai/pull/427

It adds {{gen_ai.run_guardrail.internal}} (in-process checks, which fits a 
local ONNX expert) and a {{gen_ai.security.finding}} event with:
* {{gen_ai.security.guardrail.*}}: which guardrail ran (the expert and pinned 
artifact identity);
* {{gen_ai.security.verdict.*}}: what it returned (BENIGN/INJECTION);
* {{gen_ai.security.risk.*}}: classification and score (P(INJECTION); a 
category such as OWASP LLM01 Prompt Injection);
* {{gen_ai.security.action.type}}: what the *caller* enforced, kept separate 
from the verdict;
* {{gen_ai.security.policy.*}} / {{policy.rule.id}}: which policy fired.

The verdict/action split matches this issue's rule that the model returns a 
verdict while block / quarantine / review / continue belongs to Camel routing 
policy. It also fits the requirement not to log the submitted text: none of 
these attributes needs the input. Worth keeping in mind for the capability 
descriptor and result shape, so the expert can emit these once #427 lands.

_Claude Code on behalf of davsclaus_

> camel-wolf-defender: Add a semantic expert for prompt-injection detection
> -------------------------------------------------------------------------
>
>                 Key: CAMEL-25311
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25311
>             Project: Camel
>          Issue Type: New Feature
>          Components: camel-ai
>            Reporter: Luigi De Masi
>            Assignee: Luigi De Masi
>            Priority: Major
>
> h2. Problem and intended behaviour
> Add a dedicated Wolf-Defender semantic expert in 
> {{components/camel-ai/camel-wolf-defender}}, published as 
> {{org.apache.camel:camel-wolf-defender}}, so Camel routes can screen 
> untrusted text for prompt injection through {{camel-semantic}}.
> For example, a route receiving a retrieved document or tool response should 
> be able to evaluate {{ref:injection}} and route the exchange to processing, 
> rejection, or review according to application policy. The route should not 
> need to know the model's label IDs, tokenizer, tensor format, inference 
> library, or document-scoring implementation.
> This is the concrete provider implementation associated with CAMEL-25310, 
> which introduces expert selection, capability reporting, optional 
> instructions, and catalog metadata. Reuse those contracts rather than 
> implementing a separate provider-selection mechanism in this module. Use 
> *expert* in configuration and documentation; implementing the existing 
> {{SemanticAdapter}} SPI does not require renaming that Java API.
> h2. Scope and packaging
> * Create {{camel-wolf-defender}} under the {{camel-ai}} parent and integrate 
> its dependencies, service discovery, generated metadata, documentation, and 
> build registration using Camel conventions.
> * Implement a dedicated semantic adapter/expert for the fixed 
> prompt-injection evaluation.
> * Keep {{camel-semantic}} independent of concrete inference and tokenization 
> libraries. These dependencies belong to the optional Wolf-Defender module.
> * Do not depend on LangChain4j or require a Jev-compatible service. 
> Applications should keep the common semantic contract when combining this 
> expert with other providers.
> * Target local inference for the initial implementation. ONNX Runtime's Java 
> API is a candidate backend; select and document the runtime/tokenizer 
> combination after checking the chosen export. A generic backend framework, 
> remote model-serving protocol, and separate routing endpoint are not 
> prerequisites.
> * Establish a tested baseline with a pinned Wolf-Defender v2 artifact, such 
> as Small. Document the supported model/export/runtime combinations; do not 
> imply that all variants and quantizations have been validated.
> h2. Input and capability contract
> Expose the following capabilities through the framework from CAMEL-25310:
> ||Property||Required contract||
> |Input|Text selected by the semantic declaration's state expression|
> |Result type|BOOLEAN|
> |Positive meaning|Prompt injection or jailbreak-like instruction detected|
> |Probability|Probability assigned to the INJECTION class|
> |Instructions|Unsupported; no question or synthetic instruction prompt is 
> required|
> |Caller-defined criteria|Unsupported|
> |CHOICE / SCORE|Unsupported|
> The dedicated expert supplies the meaning of the evaluation. No new {{task}} 
> keyword is needed, and users do not need to repeat a question such as "Is 
> this a prompt injection?".
> Reject CHOICE/SCORE declarations, unsupported instructions, and arbitrary 
> criteria during declaration validation, before inference. Errors should 
> identify the named evaluation and selected expert. Never ignore unsupported 
> options or silently select a different provider. Validate the selected 
> runtime value as text; structured objects require explicit application 
> selection or conversion rather than an implicit {{toString()}}.
> Preserve the exchange content and existing semantic result publication 
> behaviour. The expert's result should not replace the original message body 
> with a model-specific response.
> h2. Mapping model output to SemanticResult
> The current Wolf-Defender Small model card defines class 0 as BENIGN and 
> class 1 as INJECTION. Its ONNX example returns logits. The adapter must 
> validate the selected artifact's output contract and map the positive class 
> correctly.
> * Convert two-class logits into probabilities using numerically stable 
> softmax, where that is the selected export's contract. Do not apply softmax a 
> second time to an already normalized probability output.
> * Return {{P(INJECTION)}} through the existing BOOLEAN probability 
> representation. For example, a top-label response equivalent to BENIGN with 
> probability 0.97 must yield an injection probability of 0.03, not 0.97.
> * Validate output dimensions, label mapping, finite numeric values, and 
> probability bounds. Incompatible artifacts and malformed outputs are errors.
> * Preserve probability information until the common semantic threshold and 
> uncertainty policy are applied. Do not collapse it to the model's default 
> label first or introduce a conflicting hidden decision threshold.
> * Do not advertise this probability as an application-defined SCORE severity 
> scale, or invent a calibrated confidence guarantee.
> BENIGN means that this evaluation did not detect injection. It does not 
> establish general safety, authorization, or the absence of other threats. 
> Actions such as block, quarantine, review, or continue belong to Camel 
> routing policy; they are not additional model labels.
> h2. Tokenization and document handling
> Own tokenization, input tensors, special tokens, attention masks, padding, 
> and model invocation inside the expert. Use tokenizer assets compatible with 
> the pinned model and validate against reference inference.
> The Small v2 model card describes a 2,048-token window. Its published 
> document evaluation uses overlapping windows with 64-token overlap and 
> normalized Smooth-Max aggregation. A single truncated window does not 
> reproduce that protocol.
> Define and document bounded handling of longer input. Either implement a 
> documented document-scoring policy or reject inputs beyond the supported 
> limit explicitly; never silently discard an unchecked suffix. If adopting the 
> published aggregation, verify the exact formula and parameters against an 
> authoritative implementation rather than guessing from its name. Clearly 
> identify any alternative policy and its implications for threshold selection.
> For supported multi-window evaluation, test injection-bearing content beyond 
> the first window and at window boundaries. Limits should bound both 
> tokenization work and the number of inference windows. Empty, missing, and 
> oversized input must have an explicit, tested outcome.
> h2. Expert configuration and lifecycle
> Keep model-specific configuration on the configured Wolf expert instance: 
> model/tokenizer locations and revision, selected export, runtime options, and 
> input/document limits. Named evaluations retain common semantic options such 
> as state, threshold, and uncertainty policy.
> Support reproducible use of explicitly provisioned local artifacts without 
> requiring downloads during normal route execution. Model weights should not 
> be bundled into the Camel source repository or ordinary module artifact. 
> Document acquisition, artifact identity, and compatible 
> tokenizer/configuration files.
> Load and reuse model resources through Camel-managed lifecycle. Separate 
> declaration/capability inspection from model loading and inference. Close 
> sessions, tensors, and tokenizer/native resources on shutdown and startup 
> failure. Support concurrent evaluations without cross-exchange state leakage, 
> and bound concurrency and resource consumption using the selected runtime's 
> facilities.
> Respect the semantic SPI's timeout, interruption, and shutdown contract. 
> Verify the actual native-runtime cancellation behaviour; a timeout around a 
> Java task must not be described as stopping native inference if it only stops 
> waiting for it.
> Model-loading failures, invalid inputs, inference failures, malformed output, 
> and cancellation must remain errors. They must not become a benign result or 
> an injection probability of zero. Diagnostics should include enough 
> provider/artifact identity to reproduce a problem without logging submitted 
> text or credentials.
> h2. Integration with experts and Camel Catalog
> Register the expert through the discovery/lifecycle mechanism agreed in 
> CAMEL-25310 and allow explicitly configured registry instances. Reuse its 
> explicit/default/sole-expert resolution rules and its error on ambiguity. Two 
> configured Wolf instances may use different artifacts or limits and must 
> retain their separate identities.
> Publish generated static capability metadata for {{wolf-defender}}, linked to 
> {{camel-wolf-defender}}, through the same authoritative definition used by 
> runtime capabilities. Catalog inspection must work without downloading 
> weights or initializing a native runtime. Static metadata describes the 
> provider's contract; it must not claim that an arbitrary local artifact has 
> been validated.
> Support the existing single-evaluation and batch SPI contracts. Framework 
> grouping across experts remains the responsibility of CAMEL-25310. The 
> initial implementation may use the existing sequential batch behaviour; 
> optimized tensor batching is not required. Preserve named result keys and 
> all-or-error publication.
> h2. Illustrative application declaration
> The following uses the syntax proposed in CAMEL-25310 and assumes a 
> configured Wolf expert instance named {{security}}. It is illustrative, 
> pending the final framework API and binding conventions.
> {noformat}
> - semantic:
>     question:
>       injection:
>         expert: security
>         type: boolean
>         state: "${body}"
>         threshold: "{{security.injection.threshold}}"
>         uncertainty: "{{security.injection.uncertainty}}"
>         uncertaintyPolicy: fail
> {noformat}
> Routes use {{ref:injection}} through the semantic language. Provide a 
> runnable example showing expert construction/configuration, artifact 
> provisioning, and routing on the resulting decision. Demonstrate that 
> model-specific options remain outside the named evaluation. Explain how 
> uncertainty and operational errors are handled by the route.
> h2. License and documentation
> The Small model card declares Apache License 2.0 and retains MIT terms for 
> the upstream mmBERT-small portions. Preserve applicable notices for any 
> redistributed material and document model provenance separately from 
> runtime/tokenizer dependency licenses. Confirm the license of the exact 
> selected artifacts as part of implementation.
> Document installation, supported runtime/platform combinations, 
> configuration, fixed BOOLEAN semantics, probability mapping, limits, and 
> error behaviour. Describe false-positive/false-negative limitations without 
> making benchmark or latency claims for an untested Java implementation.
> h2. Validation and acceptance criteria
> # The new {{components/camel-ai/camel-wolf-defender}} module builds and 
> integrates with Camel's normal generated metadata and documentation processes.
> # A named BOOLEAN evaluation without instructions works through 
> {{camel-semantic}}, with both explicit expert selection and the framework's 
> sole-expert discovery behaviour.
> # Unsupported result types, instructions, and criteria fail during 
> declaration validation without running inference. Unknown/ambiguous expert 
> selection follows CAMEL-25310.
> # Tests verify positive-class mapping, stable probability conversion, 
> malformed/non-finite output rejection, and common threshold/uncertainty 
> boundaries, including a high-confidence BENIGN result.
> # Input validation and document limits are explicit and tested. Oversized 
> input is never silently partially screened; any supported windowed policy has 
> boundary and late-window coverage.
> # Lifecycle and failure tests cover resource cleanup, concurrent use, 
> cancellation/timeout behaviour, and errors remaining distinct from negative 
> classifications.
> # Batch evaluation preserves keys, expert-instance identity, and complete 
> result validation without exposing partial success.
> # Static capabilities are discoverable through Camel Catalog, agree with 
> runtime reporting, and can be read without model/native-runtime 
> initialization.
> # Normal unit tests use deterministic fixtures and require neither network 
> access nor a large model download. Add an opt-in integration test using a 
> pinned real artifact to verify tokenizer/inference parity against reference 
> outputs with a documented numeric tolerance.
> # Provide a runnable local example and setup instructions, including the 
> tested artifact/runtime, license/provenance, and representative benign, 
> injection, and difficult-benign inputs. Keep adapter correctness evidence 
> separate from claims about model detection quality.
> h2. Related issue and references
> * [CAMEL-25310: semantic experts, capabilities and catalog 
> metadata|https://issues.apache.org/jira/browse/CAMEL-25310]
> * [Wolf-Defender Small model card and inference 
> examples|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small]
> * [Wolf-Defender full 
> model|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection]
> * [Wolf-Defender Small license 
> file|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small/blob/main/LICENSE]
> * [Wolf-Defender v2 
> article|https://patronus.studio/en/posts/wolf-defender-v2-prompt-injection-detection-on-device]
> * [ONNX Runtime Java 
> API|https://onnxruntime.ai/docs/get-started/with-java.html]



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to