Oliver Caballero Alvarez created PARQUET-2493:
-------------------------------------------------
Summary: Unnecessary loss of information when opening a
HadoopInputFile
Key: PARQUET-2493
URL: https://issues.apache.org/jira/browse/PARQUET-2493
Project: Parquet
Issue Type: Improvement
Components: parquet-hadoop
Affects Versions: 1.14.0
Reporter: Oliver Caballero Alvarez
In the current version of the HadoopInputFile implementation:
[https://github.com/apache/parquet-java/blob/apache-parquet-1.14.0/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/util/HadoopInputFile.java]
When performing a newStream, the reference to the FileStatus is lost, which has
already been previously consulted to create this class. This means that when
you go to the implementation of each FileSystem, it will surely have to be
requested again, since you have requested the reference of whether the file
exists, when the file is weighed or relevant information to be able to open the
file.
To do this, it is necessary to change the Hadoop implementation, including that
in the open function, the Path object does not necessarily have to be passed
but the FileStatus can be passed directly.
return HadoopStreams.wrap(fs.open(stat.getPath()));
to
return HadoopStreams.wrap(fs.open(stat));
The implementation change is proposed in:
https://issues.apache.org/jira/browse/HADOOP-19199
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]