LuciferYang opened a new pull request, #10214:
URL: https://github.com/apache/paimon/pull/10214

   ### Purpose
   
   `DataEvolutionSplitRead.createUnionReader` builds one reader per field bunch 
in a loop, and each reader holds open file streams from the moment it is 
constructed. When a later bunch's reader failed to build (missing or corrupt 
blob/vector file), the readers already built for the earlier bunches were never 
closed and leaked their streams for the lifetime of the query.
   
   The same gap existed one level down inside a blob bunch: 
`BlobFallbackRecordReader` opens its per-sequence-group single-file readers 
eagerly, so a later group failing left the groups built before it open.
   
   Both loops are now wrapped in try/catch that closes the already-built 
readers before rethrowing, suppressing any per-reader close failure onto the 
original exception.
   
   ### Tests
   
   Added 
`DataEvolutionSplitReadTest.testFailingLaterBunchClosesEarlierBunchReaders`. It 
sets up one bunch backed by a real ORC file (whose reader opens a stream at 
construction) and a second, blob bunch pointing at a file with no bytes on disk 
so its reader fails to build, then asserts through `TraceableFileIO` that no 
input stream under the table path remains open after `createReader` throws. The 
assertion fails against current upstream (the ORC stream stays open) and passes 
with the fix.
   
   ### API and Format
   
   No.
   
   ### Documentation
   
   No.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to