peterphitran commented on code in PR #18238:
URL: https://github.com/apache/iceberg/pull/18238#discussion_r4202606057


##########
parquet/src/main/java/org/apache/iceberg/parquet/Parquet.java:
##########
@@ -1552,7 +1556,6 @@ public <D> CloseableIterable<D> build() {
           optionsBuilder.withDecryption(fileDecryptionProperties);
         }
 
-        optionsBuilder.withUseHadoopVectoredIo(true);

Review Comment:
   Yes no way to turn it off verified this myself and you're right I can look 
to tackle this after this PR. This also ties back to #16614 where 
venkateshwaracholan was working on a solution for a read option for Spark for 
that upgrade issue, I also checked on other engines looks to also be a possible 
concern for Flink and Hive.
   
   Side note: did some digging and 
https://github.com/apache/parquet-java/pull/3726 looks to split vectored reads 
at parquet.read.allocation.size, which would address the whole range allocation 
behind the OOM in #16600. I tested it locally with Iceberg: a 12.4 MB range 
became twelve 1 MB ranges. Combined with this PR, the allocation size passed 
through set() now reaches Parquet.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to