[ 
https://issues.apache.org/jira/browse/CAMEL-25311?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Claus Ibsen resolved CAMEL-25311.
---------------------------------
    Resolution: Fixed

> camel-wolf-defender: Add a semantic expert for prompt-injection detection
> -------------------------------------------------------------------------
>
>                 Key: CAMEL-25311
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25311
>             Project: Camel
>          Issue Type: New Feature
>          Components: camel-ai
>            Reporter: Luigi De Masi
>            Assignee: Luigi De Masi
>            Priority: Major
>             Fix For: 4.23.0
>
>
> h2. Problem and intended behaviour
> Add a dedicated Wolf-Defender semantic expert in 
> {{components/camel-ai/camel-wolf-defender}}, published as 
> {{org.apache.camel:camel-wolf-defender}}, so Camel routes can screen 
> untrusted text for prompt injection through {{camel-semantic}}.
> For example, a route receiving a retrieved document or tool response should 
> be able to evaluate {{ref:injection}} and route the exchange to processing, 
> rejection, or review according to application policy. The route should not 
> need to know the model's label IDs, tokenizer, tensor format, inference 
> library, or document-scoring implementation.
> This is a concrete provider implementation using the expert selection and 
> lifecycle support from CAMEL-25310 and the expert-owned evaluation contracts 
> merged in CAMEL-25382 (PR #27494). Implement the current {{SemanticAdapter}} 
> SPI with {{SemanticEvaluation}} declarations and a static {{@SemanticExpert}} 
> / {{@SemanticOperation}} contract. Reuse the framework's selection, 
> validation, batching and result-publication mechanisms.
> h2. Scope and packaging
> * Create {{camel-wolf-defender}} under the {{camel-ai}} parent and integrate 
> its dependencies, service discovery, generated metadata, documentation, and 
> build registration using Camel conventions.
> * Implement a dedicated semantic adapter/expert for the fixed 
> prompt-injection evaluation.
> * Keep {{camel-semantic}} independent of concrete inference and tokenization 
> libraries. These dependencies belong to the optional Wolf-Defender module.
> * Do not depend on LangChain4j or require a Jev-compatible service. 
> Applications should keep the common semantic contract when combining this 
> expert with other providers.
> * Keep this issue scoped to local inference. Use ONNX Runtime Java and a 
> compatible tokenizer for the validated export. Remote serving and the 
> Patronus Scanner API are outside this issue; do not add a generic backend 
> framework or a separate routing endpoint.
> * Establish the baseline with Wolf-Defender Small v2 revision 
> {{bcab2eff97bcabd7227849639e2d0d7a61b46c92}}, its FP32 ONNX export, ONNX 
> Runtime Java 1.30.0 and DJL Hugging Face Tokenizers 0.38.0. Verify artifact 
> identity and reference parity. Document tested platforms; do not imply that 
> other variants, quantizations or platforms have been validated.
> h2. Input and capability contract
> Declare a static operation contract using the framework from CAMEL-25382:
> ||Property||Required contract||
> |Operation|{{injection}}|
> |Input|Nonblank text selected by the evaluation's state expression, within 
> configured character and token limits|
> |Result type|BOOLEAN|
> |Positive meaning|Prompt injection or jailbreak-like instruction detected|
> |Probability|{{P(INJECTION)}}, retained alongside the expert's Boolean 
> verdict|
> |Instructions|Unsupported; no question or synthetic instruction prompt is 
> required|
> |Caller-defined criteria|Unsupported|
> |Other operations|Unsupported, including {{boolean}}, {{choice}}, {{score}} 
> and {{classify}}|
> |Parameters|{{threshold}}, {{uncertainty}}, {{uncertaintyPolicy}}; meanings 
> and omission behaviour are declared by this expert|
> The operation supplies the meaning of the evaluation. Applications name an 
> evaluation and select {{operation: injection}}; they do not supply a 
> question, result type or classification taxonomy.
> Use the shared static-contract validation to reject unknown operations, 
> instructions, criteria and other unsupported parameters before inference. 
> Validate expert-specific parameter combinations as well. Errors should 
> identify the named evaluation and selected expert. Never ignore unsupported 
> options or silently select a different provider. Implement {{validateInput}} 
> so every selected expert's message-dependent input requirements, including 
> token limits, are checked before any inference in a batch. Structured objects 
> require explicit application selection or conversion rather than an implicit 
> {{toString()}}.
> Preserve the exchange content and existing semantic result publication 
> behaviour. The expert's result should not replace the original message body 
> with a model-specific response.
> h2. Mapping model output to SemanticResult
> The current Wolf-Defender Small model card defines class 0 as BENIGN and 
> class 1 as INJECTION. Its ONNX example returns logits. The adapter must 
> validate the selected artifact's output contract and map the positive class 
> correctly.
> * Convert two-class logits into probabilities using numerically stable 
> softmax, where that is the selected export's contract. Do not apply softmax a 
> second time to an already normalized probability output.
> * Return a typed Boolean {{SemanticResult.value}} and retain {{P(INJECTION)}} 
> in {{SemanticResult.probability}}. For example, BENIGN probability 0.97 in 
> this exhaustive two-class model corresponds to injection probability 0.03, 
> not 0.97. Do not return a probability-only Boolean result.
> * Validate output dimensions, label mapping, finite numeric values, and 
> probability bounds. Incompatible artifacts and malformed outputs are errors.
> * The expert applies decision policy exactly once, after probability 
> conversion. Camel does not derive another Boolean from the probability or 
> insert policy defaults. Do not collapse to the model's default label first.
> * Declare {{threshold}} as an inclusive probability cutoff in [0,1], omitted 
> meaning 0.5. Declare {{uncertainty}} as an inclusive band half-width, omitted 
> meaning 0; reject combinations whose band extends outside [0,1]. Declare 
> {{uncertaintyPolicy}} as {{fail}} (the omission behaviour) or explicit 
> {{non-match}}. With a positive band, {{fail}} raises an error inside it; 
> {{non-match}} returns false while retaining the probability. These are 
> Wolf-Defender operation parameters, not universal semantic-language policies.
> * Do not advertise this probability as an application-defined SCORE severity 
> scale, or invent a calibrated confidence guarantee.
> BENIGN means that this evaluation did not detect injection. It does not 
> establish general safety, authorization, or the absence of other threats. 
> Actions such as block, quarantine, review, or continue belong to Camel 
> routing policy; they are not additional model labels.
> h2. Tokenization and document handling
> Own tokenization, input tensors, special tokens, attention masks, padding, 
> and model invocation inside the expert. Use tokenizer assets compatible with 
> the pinned model and validate against reference inference.
> The Small v2 model card describes a 2,048-token window. Its published 
> document evaluation uses overlapping windows with 64-token overlap and 
> normalized Smooth-Max aggregation. A single truncated window does not 
> reproduce that protocol.
> Define and document bounded handling of longer input. Either implement a 
> documented document-scoring policy or reject inputs beyond the supported 
> limit explicitly; never silently discard an unchecked suffix. If adopting the 
> published aggregation, verify the exact formula and parameters against an 
> authoritative implementation rather than guessing from its name. Clearly 
> identify any alternative policy and its implications for threshold selection.
> For supported multi-window evaluation, test injection-bearing content beyond 
> the first window and at window boundaries. Limits should bound both 
> tokenization work and the number of inference windows. Empty, missing, and 
> oversized input must have an explicit, tested outcome.
> h2. Expert configuration and lifecycle
> Keep model-specific configuration on the configured Wolf expert instance: 
> local artifact directory, runtime options, and input/document limits. Named 
> evaluations contain {{expert}}, {{operation}}, {{state}} and expert-owned 
> {{parameters}}. Configured instances cannot redefine the operation contract 
> or parameter semantics.
> Support reproducible use of explicitly provisioned local artifacts without 
> requiring downloads during normal route execution. Validate the pinned model, 
> tokenizer and configuration hashes. Model weights must not be bundled into 
> the Camel source repository or ordinary module artifact. Document 
> acquisition, artifact identity, compatible files, and native-library offline 
> setup.
> Load and reuse model resources through Camel-managed lifecycle. Separate 
> declaration/capability inspection from model loading and inference. Close 
> sessions, tensors, and tokenizer/native resources on shutdown and startup 
> failure. Support concurrent evaluations without cross-exchange state leakage, 
> and bound concurrency and resource consumption using the selected runtime's 
> facilities.
> Respect the semantic SPI's timeout, interruption, and shutdown contract. 
> Verify the actual native-runtime cancellation behaviour; a timeout around a 
> Java task must not be described as stopping native inference if it only stops 
> waiting for it. A timed-out worker retains its resources until it exits; 
> reject restart while that worker is still running.
> Model-loading failures, invalid inputs, inference failures, malformed output, 
> and cancellation must remain errors. They must not become a benign result or 
> an injection probability of zero. Diagnostics should include enough 
> provider/artifact identity to reproduce a problem without logging submitted 
> text or credentials.
> h2. Expert discovery and module metadata
> Register the expert through the discovery/lifecycle mechanism agreed in 
> CAMEL-25310 and allow explicitly configured registry instances. Reuse its 
> explicit/default/sole-expert resolution rules and its error on ambiguity. Two 
> configured Wolf instances may use different artifacts or limits and must 
> retain their separate identities.
> Expose the static operation contract through 
> {{SemanticCapabilities.from(WolfDefenderSemanticAdapter.class)}} without 
> constructing the expert, downloading weights or initializing native 
> resources. Generate the normal module metadata and catalog/documentation 
> entries for {{wolf-defender}}. Generated expert-operation descriptors, 
> expert-specific offline validation and expert Catalog API extensions were 
> explicitly left as separate framework work by CAMEL-25382 and are not 
> acceptance criteria for this provider.
> Support the existing single-evaluation and batch SPI contracts. Framework 
> grouping across experts remains the responsibility of CAMEL-25310. The 
> initial implementation may use the existing sequential batch behaviour; 
> optimized tensor batching is not required. Preserve named result keys and 
> all-or-error publication.
> h2. Illustrative application declaration
> The following uses the merged CAMEL-25382 declaration model and assumes a 
> configured Wolf expert instance named {{security}}:
> {noformat}
> - semantic:
>     evaluation:
>       injection:
>         expert: security
>         operation: injection
>         state: "${body}"
>         parameters:
>           threshold: !number "{{security.injection.threshold}}"
>           uncertainty: !number "{{security.injection.uncertainty}}"
>           uncertaintyPolicy: fail
> {noformat}
> Routes can invoke {{${semantic('injection')}}} through Simple, or use 
> {{ref:injection}} through the semantic language. Provide a runnable example 
> showing expert construction/configuration, artifact provisioning, and routing 
> on the resulting decision. Keep model-specific options outside the named 
> evaluation. Explain how uncertainty and operational errors reach the route's 
> review/error handling.
> h2. License and documentation
> The Small model card declares Apache License 2.0 and retains MIT terms for 
> the upstream mmBERT-small portions. Preserve applicable notices for any 
> redistributed material and document model provenance separately from 
> runtime/tokenizer dependency licenses. Confirm the license of the exact 
> selected artifacts as part of implementation.
> Document installation, supported runtime/platform combinations, 
> configuration, fixed BOOLEAN semantics, probability mapping, limits, and 
> error behaviour. Describe false-positive/false-negative limitations without 
> making benchmark or latency claims for an untested Java implementation.
> h2. Validation and acceptance criteria
> # The new {{components/camel-ai/camel-wolf-defender}} module builds and 
> integrates with Camel's normal generated metadata and documentation processes.
> # A named {{injection}} evaluation without instructions returns a Boolean 
> through {{camel-semantic}} and Simple, with explicit expert selection and 
> sole-expert discovery.
> # Unknown operations, unsupported parameters and invalid parameter 
> combinations fail during declaration validation without inference. 
> Unknown/ambiguous expert selection follows the shared framework.
> # Tests verify positive-class mapping, stable probability conversion, 
> malformed/non-finite output rejection, expert-owned threshold/uncertainty 
> boundaries and omission behaviour, including a high-confidence BENIGN result. 
> Verify that Camel does not apply a second threshold.
> # Input validation and document limits are explicit and tested. Oversized 
> input is never silently partially screened; any supported windowed policy has 
> boundary and late-window coverage.
> # Lifecycle and failure tests cover resource cleanup, concurrent use, 
> cancellation/timeout behaviour, restart after timed-out shutdown, and errors 
> remaining distinct from negative classifications.
> # Batch evaluation preserves keys and expert-instance identity, preflights 
> input requirements before any inference, and validates all results without 
> publishing partial success.
> # Static capabilities are readable from the expert class and agree with 
> runtime validation without model/native-runtime initialization. Standard 
> module catalog metadata and documentation are generated; new expert-specific 
> Catalog APIs are excluded.
> # Normal unit tests use deterministic fixtures and require neither network 
> access nor a large model download. Add an opt-in integration test using a 
> pinned real artifact to verify tokenizer/inference parity against reference 
> outputs with a documented numeric tolerance.
> # Provide a runnable local example and setup instructions, including the 
> tested artifact/runtime, license/provenance, and representative benign, 
> injection, and difficult-benign inputs. Keep adapter correctness evidence 
> separate from claims about model detection quality.
> h2. Related issue and references
> * [CAMEL-25310: experts and runtime capability 
> discovery|https://issues.apache.org/jira/browse/CAMEL-25310]
> * [CAMEL-25382: expert-owned evaluation 
> contracts|https://issues.apache.org/jira/browse/CAMEL-25382]
> * [Merged expert-contract implementation, PR 
> #27494|https://github.com/apache/camel/pull/27494]
> * [Wolf-Defender Small model card and inference 
> examples|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small]
> * [Wolf-Defender full 
> model|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection]
> * [Wolf-Defender Small license 
> file|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small/blob/main/LICENSE]
> * [Wolf-Defender v2 
> article|https://patronus.studio/en/posts/wolf-defender-v2-prompt-injection-detection-on-device]
> * [ONNX Runtime Java 
> API|https://onnxruntime.ai/docs/get-started/with-java.html]
> _Description updated with Codex on behalf of @luigidemasi to align with the 
> merged CAMEL-25382 contract._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to