shashank created CAMEL-25456:
--------------------------------
Summary: camel-wasm - a trap in the guest function leaks its
memory on the shared instance, and the cleanup releases the error buffer with
the wrong size or hides the original exception
Key: CAMEL-25456
URL: https://issues.apache.org/jira/browse/CAMEL-25456
Project: Camel
Issue Type: Bug
Reporter: shashank
A {{wasm}} producer and a {{wasm}} expression each create one Endive
{{Instance}} and use it for every message, one at a time under the lock of
{{WasmFunction}}. When the guest function fails without returning (a trap: a
Rust {{panic!}}/{{unwrap()}} on bad input, {{unreachable}}, out-of-bounds
access) or the thread is interrupted, {{WasmFunction.run()}} (main
e1bdb2f124b4) keeps using that instance:
# *Memory leak.* What the guest allocated before the trap stays allocated and
nothing resets the instance. With the {{functions.wasm}} module of the
camel-wasm tests ({{process}} does {{String::from_utf8(msg.body).unwrap()}}),
32 messages with a 256 KiB body that is not UTF-8 grew the linear memory from
17 to 153 pages (+8.5 MiB); a guest that allocates 1 MiB and traps grew it from
17 to 1055 pages in 64 calls. Rust and Go modules declare no memory maximum
(65536 pages = 4 GiB), so a stream of bad messages grows one endpoint's heap
until the JVM runs out of memory.
# *Error buffer released with the wrong size.* When the function sets the error
bit on the returned size, the message is read with the size without the bit,
but the {{finally}} block calls {{dealloc(ptr, size | 1 << 31)}} (line 89): the
guest received -2147483644 instead of 4. With the {{dealloc}} of the
documentation ({{Vec::from_raw_parts(ptr, 0, len)}}) that is a capacity above
{{isize::MAX}}, undefined behaviour.
# *Original exception replaced.* After a failure the {{finally}} block calls
{{dealloc}} on the failed instance (line 86). On an interrupted thread that
call throws again (Endive's interruption check reads the flag without clearing
it), so the {{WasmInterruptedException}} comes from the {{dealloc}}, the
exception of the call is lost and the input buffer is never released. Any other
failure of that cleanup call replaces the original exception the same way.
h3. Reproduction
A test guest written in WAT (compiled at test time with {{run.endive:wabt}}):
one page of memory, an arena allocator that only gets its memory back when
every allocation was released, a {{dealloc}} that traps on a size with the
error bit set. On main (two runs):
* 3 calls that allocate 32 KiB and trap, then a 16 KiB echo: {{TrapException:
Trapped on unreachable instruction}} (the memory of the failed calls was never
released);
* a call that returns the error {{boom}}: {{TrapException}} from {{dealloc}}
instead of {{RuntimeException: boom}};
* a call interrupted while the guest runs (40 KiB input), then a 16 KiB echo:
{{TrapException}} (the input buffer of the interrupted call was never released).
A TLA+ model of {{run()}} (exchanges sharing one instance, non-atomic guest
allocator, shared operand stack, abort while the guest runs, sticky interrupt
flag) confirms the lock is correct (no cross-talk, no double allocation,
termination), that an abort leaks a block whether the interrupt flag is sticky
or cleared, and that discarding the aborted instance without calling into it
leaks nothing (38,248 states).
h3. Proposed fix
When the call does not complete (any {{RuntimeException}}/{{Error}} from alloc,
the write, the function, the read or dealloc), make no further call into the
instance, discard it and rethrow the original exception; create a new instance
from the parsed module on the next call (under the same lock, about 0.1-0.2 ms
for the 2.2 MB test module). Release the error buffer with the size without the
error bit. A normal return, also with the error bit, keeps the instance. The
rebuild is automatic: after a trap the instance is left as the failed call left
it and the host cannot know what it allocated. Visible effect, documented in
the component/language pages and the upgrade guide: state that a module keeps
in its memory or globals is reset after a failed call. Alternative: an option
to keep the instance after a failure (default off), not proposed.
Affected: since 4.4 (CAMEL-20336), main included. Under the security model
(message senders untrusted, routes trusted) this is resource exhaustion, so a
robustness bug, not a vulnerability report.
Duplicate check (2026-10-07): no camel-wasm JIRA component; text search "wasm"
with trap/dealloc finds nothing related; the only camel-wasm issues are
CAMEL-20336, CAMEL-21569 and CAMEL-24059 (all resolved).
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)