[
https://issues.apache.org/jira/browse/CAMEL-25456?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
shashank reassigned CAMEL-25456:
--------------------------------
Assignee: shashank
> camel-wasm - a trap in the guest function leaks its memory on the shared
> instance, and the cleanup releases the error buffer with the wrong size or
> hides the original exception
> --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25456
> URL: https://issues.apache.org/jira/browse/CAMEL-25456
> Project: Camel
> Issue Type: Bug
> Reporter: shashank
> Assignee: shashank
> Priority: Major
>
> A {{wasm}} producer and a {{wasm}} expression each create one Endive
> {{Instance}} and use it for every message, one at a time under the lock of
> {{WasmFunction}}. When the guest function fails without returning (a trap: a
> Rust {{panic!}}/{{unwrap()}} on bad input, {{unreachable}}, out-of-bounds
> access) or the thread is interrupted, {{WasmFunction.run()}} (main
> e1bdb2f124b4) keeps using that instance:
> # *Memory leak.* What the guest allocated before the trap stays allocated and
> nothing resets the instance. With the {{functions.wasm}} module of the
> camel-wasm tests ({{process}} does {{String::from_utf8(msg.body).unwrap()}}),
> 32 messages with a 256 KiB body that is not UTF-8 grew the linear memory from
> 17 to 153 pages (+8.5 MiB); a guest that allocates 1 MiB and traps grew it
> from 17 to 1055 pages in 64 calls. Rust and Go modules declare no memory
> maximum (65536 pages = 4 GiB), so a stream of bad messages grows one
> endpoint's heap until the JVM runs out of memory.
> # *Error buffer released with the wrong size.* When the function sets the
> error bit on the returned size, the message is read with the size without the
> bit, but the {{finally}} block calls {{dealloc(ptr, size | 1 << 31)}} (line
> 89): the guest received -2147483644 instead of 4. With the {{dealloc}} of the
> documentation ({{Vec::from_raw_parts(ptr, 0, len)}}) that is a capacity above
> {{isize::MAX}}, undefined behaviour.
> # *Original exception replaced.* After a failure the {{finally}} block calls
> {{dealloc}} on the failed instance (line 86). On an interrupted thread that
> call throws again (Endive's interruption check reads the flag without
> clearing it), so the {{WasmInterruptedException}} comes from the {{dealloc}},
> the exception of the call is lost and the input buffer is never released. Any
> other failure of that cleanup call replaces the original exception the same
> way.
> h3. Reproduction
> A test guest written in WAT (compiled at test time with {{run.endive:wabt}}):
> one page of memory, an arena allocator that only gets its memory back when
> every allocation was released, a {{dealloc}} that traps on a size with the
> error bit set. On main (two runs):
> * 3 calls that allocate 32 KiB and trap, then a 16 KiB echo: {{TrapException:
> Trapped on unreachable instruction}} (the memory of the failed calls was
> never released);
> * a call that returns the error {{boom}}: {{TrapException}} from {{dealloc}}
> instead of {{RuntimeException: boom}};
> * a call interrupted while the guest runs (40 KiB input), then a 16 KiB echo:
> {{TrapException}} (the input buffer of the interrupted call was never
> released).
> A TLA+ model of {{run()}} (exchanges sharing one instance, non-atomic guest
> allocator, shared operand stack, abort while the guest runs, sticky interrupt
> flag) confirms the lock is correct (no cross-talk, no double allocation,
> termination), that an abort leaks a block whether the interrupt flag is
> sticky or cleared, and that discarding the aborted instance without calling
> into it leaks nothing (38,248 states).
> h3. Proposed fix
> When the call does not complete (any {{RuntimeException}}/{{Error}} from
> alloc, the write, the function, the read or dealloc), make no further call
> into the instance, discard it and rethrow the original exception; create a
> new instance from the parsed module on the next call (under the same lock,
> about 0.1-0.2 ms for the 2.2 MB test module). Release the error buffer with
> the size without the error bit. A normal return, also with the error bit,
> keeps the instance. The rebuild is automatic: after a trap the instance is
> left as the failed call left it and the host cannot know what it allocated.
> Visible effect, documented in the component/language pages and the upgrade
> guide: state that a module keeps in its memory or globals is reset after a
> failed call. Alternative: an option to keep the instance after a failure
> (default off), not proposed.
> Affected: since 4.4 (CAMEL-20336), main included. Under the security model
> (message senders untrusted, routes trusted) this is resource exhaustion, so a
> robustness bug, not a vulnerability report.
> Duplicate check (2026-10-07): no camel-wasm JIRA component; text search
> "wasm" with trap/dealloc finds nothing related; the only camel-wasm issues
> are CAMEL-20336, CAMEL-21569 and CAMEL-24059 (all resolved).
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)