raghavyadav01 opened a new pull request, #19736:
URL: https://github.com/apache/pinot/pull/19736
An OPEN_STRUCT column is split at segment build time into materialized
per-key child columns plus one shared blob column holding every key that was
not materialized. Both halves were configurable in their own idiosyncratic
way rather than the way the rest of Pinot is configured. Two fixes.
## 1. A key can ask for its dictionary under `indexes`
A dictionary is enabled on an ordinary column by putting it in `indexes`. An
OPEN_STRUCT key could not be: `FieldIndexConfigsUtil#fromFieldConfig`, which
builds a key's index configs, skipped `DICTIONARY_ID` outright and took the
answer from `encodingType` alone — and the validator then rejected an
`indexes.dictionary` entry that did not already agree with it. The same JSON
meant different things depending on which kind of column you wrote it on, and
on a key it failed silently: you got a raw key and no diagnostic.
An enabled inverted index hid half of this, because it forces a dictionary on
its own. A key whose only index is a range one had nothing to rescue it — the
dictionary entry was ignored, the key stayed raw, and the range index was
then
asked for over a dictionary that did not exist.
`indexes.dictionary` can only turn a dictionary *on*. Turning one off stays
with `encodingType`, so a `disabled` entry that contradicts the column's own
encoding is still a contradiction and is still refused rather than quietly
winning. The forward index now follows the dictionary decision rather than
the
declared encoding, for the same reason it already followed `encodingType`: a
key whose dictionary was enabled this way needs a dict-encoded forward index,
or a reload rebuilds one over a dictionary that is not there.
## 2. The sparse blob column can carry index settings
The per-key children each accept a `FieldConfig`. The blob column could not —
index loading skipped it before it reached the index-config builder, so the
only way to evaluate a predicate against a key living inside the blob was a
full scan over every document.
`OpenStructIndexConfig` gains an optional `sparseFieldConfig`. It is a plain
`FieldConfig`, shaped exactly like the one you would write for any other
column:
```json
"openStruct": {
"sparseFieldConfig": {
"name": "props$__sparse__",
"indexes": { "json": {} }
}
}
```
The blob is forced to `RAW` whatever the config asks for. Each value is a
serialized document for one row, so a dictionary over it would be a
dictionary
of whole documents — it dedupes nothing and costs a round trip per lookup.
The existing `sparseJsonIndex: true` flag keeps working and is now defined as
exactly `indexes.json = {}`, so the two spellings converge on one code path
instead of being maintained separately.
## Testing
`OpenStructPerKeyIndexReloadTest` is now 9 tests. The one that fails when the
first production change alone is reverted is
`testARangeOnlyKeyCanAskForADictionaryUnderIndexes`. Two new cases build a
JSON index on the blob, via `sparseJsonIndex` and via `sparseFieldConfig`
naming the blob column directly.
369 tests green across the OPEN_STRUCT, config and loader suites; existing
per-key coverage (dictionary, range, inverted, bloom) unchanged.
## Compatibility
Additive. A table config with no `sparseFieldConfig` behaves exactly as
before, `sparseJsonIndex` produces the index it always did, and the previous
10-arg `OpenStructIndexConfig` creator is deprecated and delegating, so
existing configs deserialize unchanged.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]