jerolba opened a new issue, #3780:
URL: https://github.com/apache/parquet-java/issues/3780

   ### Describe the enhancement requested
   
   Decouple `ParquetReadOptions` from the legacy Hadoop 
`ParquetInputFormat`/`FileInputFormat` classes, which currently force pulling 
in the 
[`hadoop-mapreduce-client-core`](https://mvnrepository.com/artifact/org.apache.hadoop/hadoop-mapreduce-client-core)
 dependency (and its transitive JARs).
   
   
   ## Summary
   
   To instantiate a `org.apache.parquet.hadoop.ParquetReader` we need to use 
`org.apache.parquet.ParquetReadOptions`, which references a set of keys located 
in `org.apache.parquet.hadoop.ParquetInputFormat` and calls a static 
`getFilter` method declared also on `ParquetInputFormat`, which `extends 
org.apache.hadoop.mapreduce.lib.input.FileInputFormat`. 
   
   `ParquetInputFormat` is part of the `parquet-hadoop` module, while 
`FileInputFormat` is declared in the `hadoop-mapreduce-client-core` JAR from 
the Hadoop project.
   
   Because `ParquetReader` (via `ParquetReadOptions`) needs to call 
`ParquetInputFormat.getFilter(...)`, the JVM is forced to initialize 
`ParquetInputFormat`, and initializing a class triggers the loading and 
initialization of its superclass (`FileInputFormat`) along with its entire 
transitive dependency graph (`org.apache.hadoop.mapreduce.*`). For code that 
only needs to read a plain Parquet file (and not the MapReduce input-format 
machinery), this pulls an unwanted, heavy Hadoop-`mapreduce` dependency into 
the classpath and link set. `AvroParquetReader` and `ProtoParquetReader` extend 
from `ParquetReader` and have the same issue.
   
   `ParquetInputFormat` has three responsibilities:
   * define a set of property keys as constants
   * deserialize the filter predicates from a configuration value
   * support the integration of Parquet files into Hadoop MapReduce
   
   This issue proposes to extract the first two responsibilities into two new 
Hadoop-agnostic classes, in the `org.apache.parquet.conf` package:
   * `ParquetInputProperties`: the property-key constants only
   * `ParquetInputFilters`: the filter deserialization logic
   so that the configuration can be used without ever loading `FileInputFormat` 
and its transitive dependencies.
   
   ```
   ParquetReader                                   ← org.apache.parquet.hadoop 
(parquet-hadoop)
           │  uses
           ▼
   ParquetReadOptions                              ← org.apache.parquet 
(parquet-hadoop)
           │  (static import keys + getFilter call)
           ▼
   ParquetInputFormat.getFilter(...)               ← org.apache.parquet.hadoop 
(parquet-hadoop)
           │  extends
           ▼
   FileInputFormat<Void, T>                        ← 
org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
           │  extends
           ▼
   InputFormat<K, V>                               ← 
org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
           │  (transitive)
           ▼
   { InputSplit, JobContext, TaskAttemptContext,
     RecordReader, ... }                           ← 
org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
   ```
   
   
   With the new `ParquetInputProperties` / `ParquetInputFilters`, building 
`ParquetReadOptions` no longer touches any `org.apache.hadoop.mapreduce` type, 
so consumers not related to Hadoop avoid transitively including the MapReduce 
dependency.
   
   ## Public API / Behavioral change
   
   No behavior change. This is a refactoring:
   
   - New classes `org.apache.parquet.conf.ParquetInputProperties` and
     `org.apache.parquet.conf.ParquetInputFilters` (in `parquet-hadoop`) hold 
the constants and the
     `ParquetConfiguration`-based filter resolution respectively.
   - The legacy `org.apache.hadoop.conf.Configuration`-based entry points 
remain on
     `ParquetInputFormat` (they are used only via the legacy MapReduce path) 
and are kept for binary /
     source compatibility.
   - All pre-existing constants on `ParquetInputFormat` are now `@Deprecated` 
and delegate to the new
     classes; source and binary compatibility are preserved.
   
   ## Acceptance criteria
   
   - `ParquetReadOptions` (used for plain file reads) references only 
`ParquetInputProperties` / `ParquetInputFilters`, never `ParquetInputFormat` / 
`org.apache.hadoop.mapreduce.InputFormat`.
   - Loading `ParquetInputProperties` / `ParquetInputFilters` / 
`ParquetReadOptions` does not initialize `FileInputFormat`.
   - All existing tests still pass (`./mvnw test`).
   
   
   ### Component(s)
   
   Core


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to