LuciferYang opened a new issue, #10243:
URL: https://github.com/apache/paimon/issues/10243

   ### Search before asking
   
   - [X] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   ### Paimon version
   
   master
   
   ### Compute Engine
   
   Flink / Spark (write)
   
   ### Minimal reproduce step
   
   `SingleFileWriter` opens the data-file output stream and creates the 
UUID-named file before the `RowDataFileWriter` body runs. If the body then 
throws, the opened stream is never closed and the file is left on storage. A 
concrete trigger is `DataFileIndexWriter.create` parsing a file-index option 
value, for example `file-index.bloom-filter.<col>.items`, whose value is read 
via `Integer.parseInt` in `BloomFilterFileIndex`. 
`SchemaValidation.validateFileIndex` only checks column existence and data-type 
support, so a malformed option value passes `CREATE TABLE` and first fails at 
write time.
   
   ### What doesn't meet your expectations?
   
   The failure is not recoverable by the caller. 
`RollingFileWriterImpl.openCurrentWriter` assigns `currentWriter` only after 
the constructor returns, so when the constructor throws, `currentWriter` stays 
null and `RollingFileWriterImpl.abort` has nothing to clean up. Each failed 
write attempt leaks one open stream and one orphan data file, and these 
accumulate across retries and job restarts. A construction failure should close 
the stream and remove the orphan file.
   
   ### Anything else?
   
   The same pattern exists in `KeyValueDataFileWriter`, which also builds 
`DataFileIndexWriter` right after its super constructor. That can be a 
follow-up.
   
   ### Are you willing to submit a PR?
   
   - [X] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to