jianzhenwu opened a new issue, #13120: URL: https://github.com/apache/gluten/issues/13120
### Backend VL (Velox) ### Bug description **Expected behavior** Sort-shuffle row IDs should round-trip the page number and row offset exactly. A row offset that cannot be represented by the compact ID must be rejected or handled by moving to a new page before encoding; it must not silently resolve to a different address. **Actual behavior** The sort-shuffle compact row ID reserves 27 bits for the page offset, but the writer can allocate a page larger than 128 MiB for a large row and continue appending records. When a later row starts at an offset >= 2^27, `toCompactRowId` ORs the full offset into the packed ID while `extractPageNumberAndOffset` keeps only the lower 27 offset bits. The ID therefore decodes to the wrong address. The writer reads payload bytes at that address as the row-size header and may read past the page, causing a native crash during shuffle serialization/compression. **Observed crash evidence (Linux x86-64, Gluten Velox backend)** - A page had size/capacity 201,326,496 bytes (~192 MiB). - Sequentially parsing row headers from page start identified a valid record boundary at offset 155,697,476. The row-size header there was 1,388 bytes. - Encoding page 17 and offset 155,697,476 produces compact row ID `0x8947c144`. Decoding that ID yields page 17 and offset 21,479,748, exactly 128 MiB earlier and inside the first large row. - At the incorrectly decoded offset, the bytes were interpreted as row-size header `0xc7b6da00`, yielding a claimed record length of 3,350,649,348 bytes, far beyond the page capacity. - The executor core backtrace showed `VeloxSortShuffleWriter::stop` / `evictPartition` through `ShuffleCompressedOutputStream::Write` and `ZSTD_compressStream2` to `__memmove_avx512_unaligned_erms`, where SIGSEGV occurred. The production artifacts establish the bad-offset/read-past-page path. A standalone reproducer has not yet been isolated. The exact Gluten revision of the runtime shared library is not confirmed; the affected encoding and missing offset-range guard are present in the current `main` source inspected at commit `60e9cb77e7d921f00dced5297ec19866d45684dd`. **Suggested fix** Ensure every encoded row start fits in 27 bits. Roll over/spill to a new page before the cursor reaches `2^27`, including between rows in a batch, or redesign the compact ID so the full offset is represented. Add a defensive page-range check before reading the row header and regression coverage for a page larger than 128 MiB followed by additional rows. Masking the offset alone is not sufficient because it silently aliases distinct positions. ### Gluten version Main branch source contains the affected encoding (commit `60e9cb77e7d921f00dced5297ec19866d45684dd`). Exact runtime build/revision: not confirmed. ### Spark version Not confirmed from the available crash artifacts. ### Spark configurations Not available in the sanitized crash evidence. ### System information Linux x86-64 executor; exact runtime JDK and Velox revision not confirmed. ### Relevant logs Sanitized GDB/core evidence is summarized above. Internal application IDs, hostnames, container paths, and process addresses are intentionally omitted. This issue description was prepared with assistance from OpenAI ChatGPT. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
