[
https://issues.apache.org/jira/browse/CAMEL-25311?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18125065#comment-18125065
]
Claus Ibsen commented on CAMEL-25311:
-------------------------------------
Follow-up merged via https://github.com/apache/camel/pull/27577 (commit
7055fc016995) for 4.23.0: model integration tests now skip unsupported native
platforms.
_Claude Code on behalf of davsclaus_
> camel-wolf-defender: Add a semantic expert for prompt-injection detection
> -------------------------------------------------------------------------
>
> Key: CAMEL-25311
> URL: https://issues.apache.org/jira/browse/CAMEL-25311
> Project: Camel
> Issue Type: New Feature
> Components: camel-ai
> Reporter: Luigi De Masi
> Assignee: Luigi De Masi
> Priority: Major
> Fix For: 4.23.0
>
>
> h2. Problem and intended behaviour
> Add a dedicated Wolf-Defender semantic expert in
> {{components/camel-ai/camel-wolf-defender}}, published as
> {{org.apache.camel:camel-wolf-defender}}, so Camel routes can screen
> untrusted text for prompt injection through {{camel-semantic}}.
> For example, a route receiving a retrieved document or tool response should
> be able to evaluate {{ref:injection}} and route the exchange to processing,
> rejection, or review according to application policy. The route should not
> need to know the model's label IDs, tokenizer, tensor format, inference
> library, or document-scoring implementation.
> This is a concrete provider implementation using the expert selection and
> lifecycle support from CAMEL-25310 and the expert-owned evaluation contracts
> merged in CAMEL-25382 (PR #27494). Implement the current {{SemanticAdapter}}
> SPI with {{SemanticEvaluation}} declarations and a static {{@SemanticExpert}}
> / {{@SemanticOperation}} contract. Reuse the framework's selection,
> validation, batching and result-publication mechanisms.
> h2. Scope and packaging
> * Create {{camel-wolf-defender}} under the {{camel-ai}} parent and integrate
> its dependencies, service discovery, generated metadata, documentation, and
> build registration using Camel conventions.
> * Implement a dedicated semantic adapter/expert for the fixed
> prompt-injection evaluation.
> * Keep {{camel-semantic}} independent of concrete inference and tokenization
> libraries. These dependencies belong to the optional Wolf-Defender module.
> * Do not depend on LangChain4j or require a Jev-compatible service.
> Applications should keep the common semantic contract when combining this
> expert with other providers.
> * Keep this issue scoped to local inference. Use ONNX Runtime Java and a
> compatible tokenizer for the validated export. Remote serving and the
> Patronus Scanner API are outside this issue; do not add a generic backend
> framework or a separate routing endpoint.
> * Establish the baseline with Wolf-Defender Small v2 revision
> {{bcab2eff97bcabd7227849639e2d0d7a61b46c92}}, its FP32 ONNX export, ONNX
> Runtime Java 1.30.0 and DJL Hugging Face Tokenizers 0.38.0. Verify artifact
> identity and reference parity. Document tested platforms; do not imply that
> other variants, quantizations or platforms have been validated.
> h2. Input and capability contract
> Declare a static operation contract using the framework from CAMEL-25382:
> ||Property||Required contract||
> |Operation|{{injection}}|
> |Input|Nonblank text selected by the evaluation's state expression, within
> configured character and token limits|
> |Result type|BOOLEAN|
> |Positive meaning|Prompt injection or jailbreak-like instruction detected|
> |Probability|{{P(INJECTION)}}, retained alongside the expert's Boolean
> verdict|
> |Instructions|Unsupported; no question or synthetic instruction prompt is
> required|
> |Caller-defined criteria|Unsupported|
> |Other operations|Unsupported, including {{boolean}}, {{choice}}, {{score}}
> and {{classify}}|
> |Parameters|{{threshold}}, {{uncertainty}}, {{uncertaintyPolicy}}; meanings
> and omission behaviour are declared by this expert|
> The operation supplies the meaning of the evaluation. Applications name an
> evaluation and select {{operation: injection}}; they do not supply a
> question, result type or classification taxonomy.
> Use the shared static-contract validation to reject unknown operations,
> instructions, criteria and other unsupported parameters before inference.
> Validate expert-specific parameter combinations as well. Errors should
> identify the named evaluation and selected expert. Never ignore unsupported
> options or silently select a different provider. Implement {{validateInput}}
> so every selected expert's message-dependent input requirements, including
> token limits, are checked before any inference in a batch. Structured objects
> require explicit application selection or conversion rather than an implicit
> {{toString()}}.
> Preserve the exchange content and existing semantic result publication
> behaviour. The expert's result should not replace the original message body
> with a model-specific response.
> h2. Mapping model output to SemanticResult
> The current Wolf-Defender Small model card defines class 0 as BENIGN and
> class 1 as INJECTION. Its ONNX example returns logits. The adapter must
> validate the selected artifact's output contract and map the positive class
> correctly.
> * Convert two-class logits into probabilities using numerically stable
> softmax, where that is the selected export's contract. Do not apply softmax a
> second time to an already normalized probability output.
> * Return a typed Boolean {{SemanticResult.value}} and retain {{P(INJECTION)}}
> in {{SemanticResult.probability}}. For example, BENIGN probability 0.97 in
> this exhaustive two-class model corresponds to injection probability 0.03,
> not 0.97. Do not return a probability-only Boolean result.
> * Validate output dimensions, label mapping, finite numeric values, and
> probability bounds. Incompatible artifacts and malformed outputs are errors.
> * The expert applies decision policy exactly once, after probability
> conversion. Camel does not derive another Boolean from the probability or
> insert policy defaults. Do not collapse to the model's default label first.
> * Declare {{threshold}} as an inclusive probability cutoff in [0,1], omitted
> meaning 0.5. Declare {{uncertainty}} as an inclusive band half-width, omitted
> meaning 0; reject combinations whose band extends outside [0,1]. Declare
> {{uncertaintyPolicy}} as {{fail}} (the omission behaviour) or explicit
> {{non-match}}. With a positive band, {{fail}} raises an error inside it;
> {{non-match}} returns false while retaining the probability. These are
> Wolf-Defender operation parameters, not universal semantic-language policies.
> * Do not advertise this probability as an application-defined SCORE severity
> scale, or invent a calibrated confidence guarantee.
> BENIGN means that this evaluation did not detect injection. It does not
> establish general safety, authorization, or the absence of other threats.
> Actions such as block, quarantine, review, or continue belong to Camel
> routing policy; they are not additional model labels.
> h2. Tokenization and document handling
> Own tokenization, input tensors, special tokens, attention masks, padding,
> and model invocation inside the expert. Use tokenizer assets compatible with
> the pinned model and validate against reference inference.
> The Small v2 model card describes a 2,048-token window. Its published
> document evaluation uses overlapping windows with 64-token overlap and
> normalized Smooth-Max aggregation. A single truncated window does not
> reproduce that protocol.
> Define and document bounded handling of longer input. Either implement a
> documented document-scoring policy or reject inputs beyond the supported
> limit explicitly; never silently discard an unchecked suffix. If adopting the
> published aggregation, verify the exact formula and parameters against an
> authoritative implementation rather than guessing from its name. Clearly
> identify any alternative policy and its implications for threshold selection.
> For supported multi-window evaluation, test injection-bearing content beyond
> the first window and at window boundaries. Limits should bound both
> tokenization work and the number of inference windows. Empty, missing, and
> oversized input must have an explicit, tested outcome.
> h2. Expert configuration and lifecycle
> Keep model-specific configuration on the configured Wolf expert instance:
> local artifact directory, runtime options, and input/document limits. Named
> evaluations contain {{expert}}, {{operation}}, {{state}} and expert-owned
> {{parameters}}. Configured instances cannot redefine the operation contract
> or parameter semantics.
> Support reproducible use of explicitly provisioned local artifacts without
> requiring downloads during normal route execution. Validate the pinned model,
> tokenizer and configuration hashes. Model weights must not be bundled into
> the Camel source repository or ordinary module artifact. Document
> acquisition, artifact identity, compatible files, and native-library offline
> setup.
> Load and reuse model resources through Camel-managed lifecycle. Separate
> declaration/capability inspection from model loading and inference. Close
> sessions, tensors, and tokenizer/native resources on shutdown and startup
> failure. Support concurrent evaluations without cross-exchange state leakage,
> and bound concurrency and resource consumption using the selected runtime's
> facilities.
> Respect the semantic SPI's timeout, interruption, and shutdown contract.
> Verify the actual native-runtime cancellation behaviour; a timeout around a
> Java task must not be described as stopping native inference if it only stops
> waiting for it. A timed-out worker retains its resources until it exits;
> reject restart while that worker is still running.
> Model-loading failures, invalid inputs, inference failures, malformed output,
> and cancellation must remain errors. They must not become a benign result or
> an injection probability of zero. Diagnostics should include enough
> provider/artifact identity to reproduce a problem without logging submitted
> text or credentials.
> h2. Expert discovery and module metadata
> Register the expert through the discovery/lifecycle mechanism agreed in
> CAMEL-25310 and allow explicitly configured registry instances. Reuse its
> explicit/default/sole-expert resolution rules and its error on ambiguity. Two
> configured Wolf instances may use different artifacts or limits and must
> retain their separate identities.
> Expose the static operation contract through
> {{SemanticCapabilities.from(WolfDefenderSemanticAdapter.class)}} without
> constructing the expert, downloading weights or initializing native
> resources. Generate the normal module metadata and catalog/documentation
> entries for {{wolf-defender}}. Generated expert-operation descriptors,
> expert-specific offline validation and expert Catalog API extensions were
> explicitly left as separate framework work by CAMEL-25382 and are not
> acceptance criteria for this provider.
> Support the existing single-evaluation and batch SPI contracts. Framework
> grouping across experts remains the responsibility of CAMEL-25310. The
> initial implementation may use the existing sequential batch behaviour;
> optimized tensor batching is not required. Preserve named result keys and
> all-or-error publication.
> h2. Illustrative application declaration
> The following uses the merged CAMEL-25382 declaration model and assumes a
> configured Wolf expert instance named {{security}}:
> {noformat}
> - semantic:
> evaluation:
> injection:
> expert: security
> operation: injection
> state: "${body}"
> parameters:
> threshold: !number "{{security.injection.threshold}}"
> uncertainty: !number "{{security.injection.uncertainty}}"
> uncertaintyPolicy: fail
> {noformat}
> Routes can invoke {{${semantic('injection')}}} through Simple, or use
> {{ref:injection}} through the semantic language. Provide a runnable example
> showing expert construction/configuration, artifact provisioning, and routing
> on the resulting decision. Keep model-specific options outside the named
> evaluation. Explain how uncertainty and operational errors reach the route's
> review/error handling.
> h2. License and documentation
> The Small model card declares Apache License 2.0 and retains MIT terms for
> the upstream mmBERT-small portions. Preserve applicable notices for any
> redistributed material and document model provenance separately from
> runtime/tokenizer dependency licenses. Confirm the license of the exact
> selected artifacts as part of implementation.
> Document installation, supported runtime/platform combinations,
> configuration, fixed BOOLEAN semantics, probability mapping, limits, and
> error behaviour. Describe false-positive/false-negative limitations without
> making benchmark or latency claims for an untested Java implementation.
> h2. Validation and acceptance criteria
> # The new {{components/camel-ai/camel-wolf-defender}} module builds and
> integrates with Camel's normal generated metadata and documentation processes.
> # A named {{injection}} evaluation without instructions returns a Boolean
> through {{camel-semantic}} and Simple, with explicit expert selection and
> sole-expert discovery.
> # Unknown operations, unsupported parameters and invalid parameter
> combinations fail during declaration validation without inference.
> Unknown/ambiguous expert selection follows the shared framework.
> # Tests verify positive-class mapping, stable probability conversion,
> malformed/non-finite output rejection, expert-owned threshold/uncertainty
> boundaries and omission behaviour, including a high-confidence BENIGN result.
> Verify that Camel does not apply a second threshold.
> # Input validation and document limits are explicit and tested. Oversized
> input is never silently partially screened; any supported windowed policy has
> boundary and late-window coverage.
> # Lifecycle and failure tests cover resource cleanup, concurrent use,
> cancellation/timeout behaviour, restart after timed-out shutdown, and errors
> remaining distinct from negative classifications.
> # Batch evaluation preserves keys and expert-instance identity, preflights
> input requirements before any inference, and validates all results without
> publishing partial success.
> # Static capabilities are readable from the expert class and agree with
> runtime validation without model/native-runtime initialization. Standard
> module catalog metadata and documentation are generated; new expert-specific
> Catalog APIs are excluded.
> # Normal unit tests use deterministic fixtures and require neither network
> access nor a large model download. Add an opt-in integration test using a
> pinned real artifact to verify tokenizer/inference parity against reference
> outputs with a documented numeric tolerance.
> # Provide a runnable local example and setup instructions, including the
> tested artifact/runtime, license/provenance, and representative benign,
> injection, and difficult-benign inputs. Keep adapter correctness evidence
> separate from claims about model detection quality.
> h2. Related issue and references
> * [CAMEL-25310: experts and runtime capability
> discovery|https://issues.apache.org/jira/browse/CAMEL-25310]
> * [CAMEL-25382: expert-owned evaluation
> contracts|https://issues.apache.org/jira/browse/CAMEL-25382]
> * [Merged expert-contract implementation, PR
> #27494|https://github.com/apache/camel/pull/27494]
> * [Wolf-Defender Small model card and inference
> examples|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small]
> * [Wolf-Defender full
> model|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection]
> * [Wolf-Defender Small license
> file|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small/blob/main/LICENSE]
> * [Wolf-Defender v2
> article|https://patronus.studio/en/posts/wolf-defender-v2-prompt-injection-detection-on-device]
> * [ONNX Runtime Java
> API|https://onnxruntime.ai/docs/get-started/with-java.html]
> _Description updated with Codex on behalf of @luigidemasi to align with the
> merged CAMEL-25382 contract._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)