Arnaud-Nauwynck opened a new issue, #3076:
URL: https://github.com/apache/parquet-java/issues/3076
### Describe the enhancement requested
When reading some column chunks but not all, parquet is building a list of
"ConsecutivePartList", then trying to call the Hadoop api for vectorized reader
of FSDataInputStream#readVectored(List<FileRange> ...)
Unfortunatly, many implementations of "FSDataInputStream" do not override
the readVectored() method, which trigger many distinct calls to read.
For example on hadoop-azure, the Azure Datalake Storage is much slower at
establishing a new Https connection (using infamous calls HttpURLConnection for
jdk 1.0, then doing TLS hand-shake), that to get only few more megas of data on
an existing socket !!
The case with small wholes to avoid reading is very frequent when having
columns in parquet files that are not read, and are highly compressed because
of RLE encoding. Typically, a very sparse column with only few values, or even
always null within a page. Such a column could be encoded in only few hundred
of bytes by parquet, so it is NOT a problem of reading 100 bytes more.
Parquet should at least honor the following method from hadoop class
FileSystem, that says that a seek of less than 4096 bytes is NOT reasonable.
```
/**
* What is the smallest reasonable seek?
* @return the minimum number of bytes
*/
default int minSeekForVectorReads() {
return 4 * 1024;
}
```
The logic for building this List<ConsecutivePartList> for a list of column
chunks is here:
org.apache.parquet.hadoop.ParquetFileReader#internalReadRowGroup
```
private ColumnChunkPageReadStore internalReadRowGroup(int blockIndex)
throws IOException {
...
for (ColumnChunkMetaData mc : block.getColumns()) {
...
// first part or not consecutive => new list
if (currentParts == null || currentParts.endPos() != startingPos) {
// <===== SHOULD honor minSeekForVectorReads()
currentParts = new ConsecutivePartList(startingPos);
allParts.add(currentParts);
}
currentParts.addChunk(new ChunkDescriptor(columnDescriptor, mc,
startingPos, mc.getTotalSize()));
}
}
// actually read all the chunks
ChunkListBuilder builder = new ChunkListBuilder(block.getRowCount());
readAllPartsVectoredOrNormal(allParts, builder);
rowGroup.setReleaser(builder.releaser);
for (Chunk chunk : builder.build()) {
readChunkPages(chunk, block, rowGroup);
}
return rowGroup;
}
```
maybe a possible implementation could be to add fictive
"ConsecutivePartList" that are to be ignored while receiving the data, but that
would avoid having some wholes in the ranges to read.
### Component(s)
_No response_
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]