peterphitran commented on code in PR #18238:
URL: https://github.com/apache/iceberg/pull/18238#discussion_r4202606057
##########
parquet/src/main/java/org/apache/iceberg/parquet/Parquet.java:
##########
@@ -1552,7 +1556,6 @@ public <D> CloseableIterable<D> build() {
optionsBuilder.withDecryption(fileDecryptionProperties);
}
- optionsBuilder.withUseHadoopVectoredIo(true);
Review Comment:
Yes no way to turn it off verified this myself and you're right I can look
to tackle this after this PR. This also ties back to #16614 where
venkateshwaracholan was working on a solution for a read option for Spark for
that upgrade issue, I also checked on other engines looks to also be a possible
concern for Flink and Hive.
Side note: did some digging and
https://github.com/apache/parquet-java/pull/3726 looks to split vectored reads
at parquet.read.allocation.size, which would address the whole range allocation
behind the OOM in #16600. I tested it locally with Iceberg: a 12.4 MB range
became twelve 1 MB ranges. Combined with this PR, the allocation size passed
through set() now reaches Parquet.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]