costas-db opened a new issue, #3826:
URL: https://github.com/apache/parquet-java/issues/3826

   ### Problem
   
   Parquet column paths are component-based, but `ParquetFileWriter` stores 
Bloom filters in a map keyed by `String.join(".", descriptor.getPath())`.
   
   These distinct paths therefore collide:
   
   ```text
   Top-level field named `a.b`: ["a.b"]
   Nested field `b` in `a`:     ["a", "b"]
   ```
   
   When both columns have Bloom filters, the nested filter overwrites the 
top-level filter under the shared `"a.b"` key. During footer serialization, 
both columns can receive the nested column filter.
   
   ### Reproduction
   
   A test writes disjoint values (`top-*` and `nested-*`) to the two string 
columns and checks each footer Bloom filter for its own values.
   
   Current result:
   
   ```text
   Top-level field path ["a.b"] contains its own Bloom values: false
   Nested field path ["a", "b"] contains its own Bloom values: true
   ```
   
   Expected: both results are `true`.
   
   ### Affected code
   
   - `ParquetFileWriter.writeColumnChunk`
   - `ParquetFileWriter.serializeBloomFilters`
   
   Similar dot-string path handling also exists in dictionary lookup and 
row-group copying, but this issue includes a minimal Bloom-filter reproduction.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to