PDGGK opened a new issue, #18237:
URL: https://github.com/apache/iotdb/issues/18237

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/iotdb/issues) and 
found nothing similar.
   
   ### Version
   
   `master` (2.0.x). The affected code is also present in released 2.0.x 
versions.
   
   ### Describe the bug and provide the minimal reproduce step
   
   `AbstractDataTool.readCsvFile(String)` builds a `CSVParser` over `new 
InputStreamReader(new FileInputStream(path))` and returns it. The `CSVParser` 
owns that `FileInputStream`, but the CLI import code paths that call it never 
close the returned parser — it is assigned to a local inside a plain `try { ... 
}` block (no try-with-resources, no `finally`). Every imported file therefore 
leaks its file descriptor, and the early `return`s for an empty file or an 
invalid header leak it immediately, because the parser is opened before those 
checks run.
   
   Affected call sites (current `master`, also present in released 2.0.x):
   
   - 
`iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportData.java` — 
`importFromSingleFile`
   - 
`iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTree.java` 
— `importFromCsvFile`
   - 
`iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTable.java`
 — `importFromCsvFile`
   - 
`iotdb-client/cli/src/main/java/org/apache/iotdb/tool/schema/ImportSchemaTree.java`
 — `importSchemaFromCsvFile` (this class has its own copy of `readCsvFile`)
   
   Minimal reproduce step:
   
   1. Create a directory containing a large number of small CSV files — more 
than the process open-file limit (for example a few thousand files under 
`ulimit -n 1024`).
   2. Run the CLI data import over that directory.
   3. The import fails partway through with `Too many open files`. (A single 
import already leaks one descriptor; it is simply not fatal until enough 
accumulate.)
   
   ### What did you expect to see?
   
   Each `CSVParser` (and the `FileInputStream` it wraps) is closed after the 
file is processed, on every exit path — including the empty-file / 
invalid-header early returns. Importing a large directory of CSV files should 
not exhaust the process's file descriptors.
   
   ### What did you see instead?
   
   The `CSVParser` returned by `readCsvFile` is never closed, so its underlying 
`FileInputStream` stays open. Importing a directory with enough CSV files leaks 
descriptors until the import fails with `Too many open files`.
   
   ### Anything else?
   
   The record `Stream` is fully consumed inside the same block before the 
method returns, so consuming each parser in a try-with-resources (closing it on 
scope exit) is safe.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to