costas-db opened a new issue, #3832:
URL: https://github.com/apache/parquet-java/issues/3832

   ### Problem
   
   Per-column Hadoop configuration currently identifies a column with a 
flattened dot string such as `a.b`. This cannot distinguish two valid, 
different Parquet paths:
   
   ```text
   Top-level field named `a.b`: ["a.b"]
   Nested field `b` in `a`:     ["a", "b"]
   ```
   
   `ColumnConfigParser` passes the suffix to `ParquetProperties` as a string, 
and `ColumnProperty` interprets it with `ColumnPath.fromDotString`. 
Consequently, a setting intended for the top-level dotted field cannot be 
represented and may instead apply to the nested field. This is 
datatype-independent and affects per-column settings such as Bloom filters and 
statistics.
   
   ### Reproduction
   
   Use two ordinary string columns with the paths above, disable Bloom filters 
and statistics globally, and enable them only for the top-level `["a.b"]` path. 
On current `master`, the top-level column still has no Bloom filter or value 
statistics because the structured override is not recognized.
   
   ### Proposed direction
   
   Add a backward-compatible structured-path key form that encodes each path 
component independently, plus path-aware parser and builder overloads. Keep 
existing dot-string keys working for compatibility, while allowing callers to 
target dotted field names unambiguously.
   
   This changes only configuration lookup identity; it does not alter the 
Parquet file format.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to