costas-db opened a new issue, #3826:
URL: https://github.com/apache/parquet-java/issues/3826
### Problem
Parquet column paths are component-based, but `ParquetFileWriter` stores
Bloom filters in a map keyed by `String.join(".", descriptor.getPath())`.
These distinct paths therefore collide:
```text
Top-level field named `a.b`: ["a.b"]
Nested field `b` in `a`: ["a", "b"]
```
When both columns have Bloom filters, the nested filter overwrites the
top-level filter under the shared `"a.b"` key. During footer serialization,
both columns can receive the nested column filter.
### Reproduction
A test writes disjoint values (`top-*` and `nested-*`) to the two string
columns and checks each footer Bloom filter for its own values.
Current result:
```text
Top-level field path ["a.b"] contains its own Bloom values: false
Nested field path ["a", "b"] contains its own Bloom values: true
```
Expected: both results are `true`.
### Affected code
- `ParquetFileWriter.writeColumnChunk`
- `ParquetFileWriter.serializeBloomFilters`
Similar dot-string path handling also exists in dictionary lookup and
row-group copying, but this issue includes a minimal Bloom-filter reproduction.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]