LuciferYang opened a new issue, #10066:
URL: https://github.com/apache/paimon/issues/10066

   ### Search before asking
   
   - [X] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   ### Paimon version
   
   master
   
   ### Compute Engine
   
   Flink / Spark (bucketed append-only table clustering)
   
   ### Minimal reproduce step
   
   On a bucketed append-only table with clustering enabled, trigger a 
compaction whose write instance clusters more than one (partition, bucket) 
group, and give the sort buffer enough data to spill (a small 
`write-buffer-size` / small page size, or enough rows). The first group's 
clustering succeeds. The second one throws `java.io.FileNotFoundException: 
.../paimon-io-*/....channel (No such file or directory)` as soon as its sort 
buffer spills.
   
   ### What doesn't meet your expectations?
   
   `Sorter.close()` closes the `IOManager` it was handed. That `IOManager` 
belongs to the write and is shared by everything the write does, and its spill 
directories are created once, when the manager is constructed. Closing it 
deletes those directories, so after the first clustering round any later spill 
of the same write instance fails with `FileNotFoundException`.
   
   Expected: clustering a table whose write spills across more than one group 
does not fail.
   
   ### Anything else?
   
   Fix direction: `Sorter` must not close the caller-owned `IOManager`. The 
sorter's own resources are released by `buffer.clear()`, which deletes its 
spill channels through the buffer's private `SpillChannelManager`, so nothing 
sorter-owned leaks. Separately, the sorter's input `RecordReaderIterator` was 
never closed by `clusterRewrite`; close it in `Sorter.close()`.
   
   ### Are you willing to submit a PR?
   
   - [X] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to