kevinjqliu opened a new issue, #3672:
URL: https://github.com/apache/parquet-java/issues/3672

   ## Motivation
   
   `parquet-cli` still uses Avro as an internal row model in places where users 
are only trying to inspect Parquet files. Valid Parquet is not always valid 
Avro, so Parquet inspection should not require Parquet schemas or records to 
round-trip through Avro.
   
   `cat/head` already has a group-reader fallback for one class of problem: 
Avro schema conversion failures such as invalid Avro field names. That fallback 
is useful, but it is still a special case. Other inspection paths, such as 
`scan`, still use the Avro-backed data path for Parquet input.
   
   Concrete examples:
   
   - Invalid Avro field names: `original-instruction` and `column with known 
type` are valid Parquet field names but invalid Avro names. This already 
required a `cat/head` group-reader fallback.
   - Nested Parquet structures: `data/nested_lists.snappy.parquet` is a valid 
`parquet-testing` fixture, but `cat` and `scan` fail through the Avro-backed 
reader.
   - Projection/type mismatch: UUID and INT96 Parquet columns can fail when 
Avro projection rewrites the requested physical type.
   - Lost Parquet structure: shredded Variant and other newer logical types can 
have valid Parquet physical structures that Avro cannot represent without 
losing information.
   
   Avro is still part of the CLI API for conversion workflows. The goal is not 
to remove Avro; it is to stop requiring Avro for Parquet inspection.
   
   ## Proposal
   
   Use Avro only when the command or input/output format requires it. Use 
Parquet-native readers and metadata APIs for Parquet inspection and physical 
rewrites.
   
   | Command | Avro needed? | Proposed internal path |
   |---|---:|---|
   | `help`, `version` | No | CLI metadata only |
   | `meta`, `pages`, `dictionary`, `check-stats`, `column-index`, 
`column-size`, `footer`, `bloom-filter`, `size-stats`, `geospatial-stats` | No 
| `ParquetFileReader` and Parquet metadata/page/stat APIs |
   | `prune`, `trans-compression`, `masking`, `rewrite` | No | Parquet rewrite 
utilities |
   | `cat`, `head` | Only for non-Parquet inputs | Parquet -> 
`GroupReadSupport`; Avro/JSON -> Avro path |
   | `scan` | Only for non-Parquet inputs | Parquet -> `GroupReadSupport`; 
Avro/JSON -> Avro path |
   | `schema` | Yes by default | Keep Avro schema output; keep physical Parquet 
path for `--parquet` |
   | `csv-schema` | Yes | Avro schema inference |
   | `convert-csv` | Yes | CSV -> Avro records/schema -> Parquet |
   | `convert` | Yes | Avro-backed conversion path |
   | `to-avro` | Yes | Avro writer/schema path |
   
   ## Related issues
   
   - https://github.com/apache/parquet-java/issues/2836 / 
https://github.com/apache/parquet-java/pull/3332: `cat` failed because Avro 
rejected a Parquet field name. The merged fix added a Parquet group-reader 
fallback.
   - https://github.com/apache/parquet-java/issues/2835: `head` reported the 
same class of invalid Avro field-name failure.
   - https://github.com/apache/parquet-java/issues/2708: `cat` failed on 
parquet-protobuf files with modern LIST encoding via `AvroRecordConverter`.
   - https://github.com/apache/parquet-java/issues/1641: UUID values failed 
because Avro projection changed the requested physical type; skipping 
projection let the file read.
   - https://github.com/apache/parquet-java/issues/3345: INT96 data failed 
through Avro requested-schema conversion.
   - https://github.com/apache/parquet-java/issues/3477: Avro schema conversion 
dropped `typed_value` from shredded Variant schemas.
   - https://github.com/apache/parquet-java/issues/2339: `convert-csv` hit Avro 
naming rules. This should stay in the Avro-backed conversion area, not drive 
Parquet inspection behavior.
   
   Tangential: https://github.com/apache/parquet-java/issues/2657 and 
https://github.com/apache/parquet-java/issues/2630 show additional CLI coupling 
to Avro internals, but are mostly binary/API mismatch issues.
   
   ## Summary
   
   Keep Avro for Avro-facing workflows: `schema` default output, `to-avro`, 
`convert`, CSV/JSON conversion, and Avro input handling.
   
   Use Parquet-native readers for Parquet inspection and physical rewrites. 
This preserves the Avro API where intentional while making `parquet-cli` more 
robust for valid Parquet files that Avro cannot model.
   
   ### Appendix
   #### Repro examples
   
   Using `parquet-mr version 1.17.1` and fixtures from `apache/parquet-testing`:
   
   ```bash
   git clone https://github.com/apache/parquet-testing.git
   cd parquet-testing
   ```
   
   These Parquet inspection commands currently fail:
   
   ```bash
   parquet cat -n 1 data/nested_lists.snappy.parquet
   parquet scan data/nested_lists.snappy.parquet
   
   parquet cat -n 1 shredded_variant/case-001.parquet
   parquet scan shredded_variant/case-001.parquet
   
   parquet cat -n 1 data/int96_from_spark.parquet
   parquet scan data/int96_from_spark.parquet
   ```
   
   Observed failure classes:
   
   - `data/nested_lists.snappy.parquet`: Avro schema conversion succeeds, but 
`cat` and `scan` fail during Avro-backed record reading with 
`ParquetDecodingException`, followed by `GroupColumnIO.getFirst` index failure.
   - `shredded_variant/case-001.parquet`: Avro schema conversion succeeds, but 
`cat` and `scan` fail during Avro-backed record reading with 
`ParquetDecodingException`, followed by `GroupColumnIO.getLast` index failure.
   - `data/int96_from_spark.parquet`: Avro schema conversion fails before 
reading records with `Argument error: INT96 is deprecated...`, blocking Parquet 
row inspection.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to