Arnaud-Nauwynck opened a new issue, #3077:
URL: https://github.com/apache/parquet-java/issues/3077

   ### Describe the enhancement requested
   
   In hadoop-azure, there are huge performance problems when reading file in a 
too fragmented way: by reading many small file fragments even with the 
readVectored() Hadoop API, resulting in distinct Https Requests (=TCP-IP 
connection established + TLS handshake + requests).
   Internally, at lowest level, haddop azure is using class HttpURLConnection 
from jdk 1.0,  and the ReadAhead Threads do not sufficiently solve all problems.
   The hadoop azure implementation of "readVectored()" should make a compromise 
between reading extra ignored data wholes, and establishing too many https 
connections.
   
   Currently, the class AzureBlobFileSystem#open() does return a default 
inneficient imlpementation of readVectored:
   ```
     private FSDataInputStream open(final Path path,
         final Optional<OpenFileParameters> parameters) throws IOException {
   ...
         InputStream inputStream = 
getAbfsStore().openFileForRead(qualifiedPath, parameters, statistics, 
tracingContext);
         return new FSDataInputStream(inputStream);  // <== FSDataInputStream 
is not efficiently overriding readVectored() !
    }
   ```
   
   see default implementation of FSDataInpustStream.readVectored:
   ```
       public void readVectored(List<? extends FileRange> ranges, 
IntFunction<ByteBuffer> allocate) throws IOException {
           ((PositionedReadable)this.in).readVectored(ranges, allocate);
       }
   ```
   
   it calls the underlying method from class AbfsInputStream, which is not 
overriden:
   ```
       default void readVectored(List<? extends FileRange> ranges, 
IntFunction<ByteBuffer> allocate) throws IOException {
           VectoredReadUtils.readVectored(this, ranges, allocate);
       }
   ```
   
   AbfsInputStream should override this method, and accept internally to do 
less Https calls, with merged range, and ignore some returned data (wholes in 
the range). 
   
   It is like honouring the parameter of hadoop FSDataInputStream (implements 
PositionedReadable)
   ```
     /**
      * What is the smallest reasonable seek?
      * @return the minimum number of bytes
      */
     default int minSeekForVectorReads() {
       return 4 * 1024;
     }
   ```
   Even this 4096 value is very conservative, and should be redined by 
AbfsFileSystem to be 4Mo or even 8mo.
   
   ask chat gpt: "on Azure Storage, what is the speed of getting 8Mo of a page 
block, compared to the time to establish a https tls handshake ?"
   The response (untrusted from chat gpt..) says :
   HTTPS/TLS Handshake: ~100–300 ms  ... is generally slower than  downloading 
8 MB from Page Blob:  on Standard Tier: ~100–200 ms / on Premium Tier: ~30–50 ms
   
   Azure Abfsclient already setup by default a lot of Threads for Prefecth Read 
Ahead, to prefetch 4Mo of data,  but it is NOT sufficent, and less efficient 
that simply implementing correctly what is already in Hadoop API : 
readVectored(). It also have the drawback of reading tons of useless data (past 
parquet blocks), that are never used.
   
   
   ### Component(s)
   
   _No response_


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to