+1 on the file-wide flag.

>From what I can tell, whether a given lower/upper bound is tight comes down
to two parts.

The first is file-level and changes over time. A DV added later deletes
some rows, and the accuracy of the original bounds are no longer
guaranteed. The new file wide flag can help record the change.

The second is column-level and static. Variable-length columns like String
and Binary can have truncated bounds, so we cannot make them reliably
tight, and we don't have a use case that needs their tightness today.
That's a property of the column type in the schema, so it doesn't need
additional metadata for recording.

So moving the tightness boolean from per-column to per-entry (tracked_file
in v4) simplifies the bookkeeping for writers without giving up anything we
care about.

One note on upgrading the existing tables to v4. Since we want that to stay
a lightweight metadata operation with no manifest rewrite, tightness is
simply undefined for existing entries. Today the merge-on-read table does
not have tight bounds. Once position deletes and DVs are rewritten to be
colocated with the data file in a single tracked-file entry, tightness
becomes a well-defined property of that entry, and a writer at that point
can set the flag accordingly.

Thanks,
Hongyue

On Tue, Oct 6, 2026 at 12:23 AM Eduard Tudenhöfner <[email protected]>
wrote:

> The file-wide flag seems to make sense to me, so +1 for that
>
> On Mon, Oct 5, 2026 at 10:22 PM Daniel Weeks <[email protected]> wrote:
>
>> +1 to Russell's analysis here.
>>
>> While we can construct hypothetical use cases where there may be benefit,
>> it seems like those use cases are unlikely and/or narrow.
>>
>> I'd also suggest going with the file-wide flag.
>>
>> -Dan
>>
>> On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer <
>> [email protected]> wrote:
>>
>>> I've been thinking through the patterns where per-column tightness would
>>> matter.
>>>
>>>    1.
>>>
>>>    *Selectively marking variable-length columns tight*. I don't have a
>>>    use case for this yet. A file-level flag can still mean every 
>>> non-truncated
>>>    column is tight.
>>>    2.
>>>
>>>    *Selectively re-tightening a few columns, for example after a column
>>>    update*. Dan's point from the discussion applies here: a workload
>>>    whose DVs invalidate stats will also invalidate new column stats. A table
>>>    left in a state where only some stats are tight due to an update seems
>>>    unlikely.
>>>    3.
>>>
>>>    *A DV writer that can write tight stats*. This is the pattern that
>>>    matters to me. The writer repairs bounds while it deletes, and it can
>>>    repair every column it tracks. We still need an explicit flag, because
>>>    otherwise readers treat any attached DV as wide and ignore the repaired
>>>    bounds. One file-level flag covers this.
>>>    4.
>>>
>>>    *A command that restores tight stats when the DV writer cannot.*
>>>    This is worth supporting when deletes are infrequent: write the DVs,
>>>    analyze once, and use the tight bounds until the next DV. If we 
>>> frequently
>>>    write DV's then we are in the same bad situation as #2. Either way the
>>>    restore is all-or-nothing. I thought for a bit that there could be a
>>>    command that only restores stats for select columns, this could make 
>>> sense
>>>    with column update files but also feels like unecessary complexity.
>>>
>>> Patterns 3 and 4 are the realistic ones, and both only need the
>>> file-wide flag. So I'm kind of leaning in that direction now.
>>>
>>> On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]> wrote:
>>>
>>>> Hi, everyone -
>>>>
>>>> In v4, we introduced the notion of *tightness* in stats. Tightness
>>>> indicates whether the upper/lower bounds truly exist in the live rows of a
>>>> file. If the stats are tight, then engines can correctly answer many
>>>> queries purely by consulting the metadata.
>>>>
>>>> However, when a DV is attached to a file, the stats are assumed to be
>>>> non-tight, so these  metadata-only query optimizations won't work anymore.
>>>>
>>>> So we want a way to represent stats tightness even in the presence of a
>>>> DV. I have a short writeup
>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0>
>>>> of the problem and a few possible solutions. We discussed this today
>>>> at the v4 adaptive metadata tree sync, and I wanted to continue the
>>>> discussion.
>>>>
>>>> Please take a look at the doc
>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r>
>>>> and let us know what you think by leaving a comment in the doc or in this
>>>> email thread.
>>>>
>>>> Best,
>>>> Anoop
>>>>
>>>

Reply via email to