Oliver Caballero Alvarez created PARQUET-2493:
-------------------------------------------------

             Summary: Unnecessary loss of information when opening a 
HadoopInputFile
                 Key: PARQUET-2493
                 URL: https://issues.apache.org/jira/browse/PARQUET-2493
             Project: Parquet
          Issue Type: Improvement
          Components: parquet-hadoop
    Affects Versions: 1.14.0
            Reporter: Oliver Caballero Alvarez


In the current version of the HadoopInputFile implementation:

[https://github.com/apache/parquet-java/blob/apache-parquet-1.14.0/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/util/HadoopInputFile.java]

 

When performing a newStream, the reference to the FileStatus is lost, which has 
already been previously consulted to create this class. This means that when 
you go to the implementation of each FileSystem, it will surely have to be 
requested again, since you have requested the reference of whether the file 
exists, when the file is weighed or relevant information to be able to open the 
file.

To do this, it is necessary to change the Hadoop implementation, including that 
in the open function, the Path object does not necessarily have to be passed 
but the FileStatus can be passed directly.

return HadoopStreams.wrap(fs.open(stat.getPath()));

to

return HadoopStreams.wrap(fs.open(stat));

 


The implementation change is proposed in:

https://issues.apache.org/jira/browse/HADOOP-19199



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to