Thanks Anoop for the clear writeup for the problem, makes sense to me for
per-file tightness for simplicity.

The above list of use cases is an interesting food for thought, from
the analysis it sounds like most will repair all bounds and not need per
column after all, hence agree.

Thanks
Szehon

On Tue, Oct 6, 2026 at 2:58 PM Hongyue Zhang <[email protected]>
wrote:

> +1 on the file-wide flag.
>
> From what I can tell, whether a given lower/upper bound is tight comes
> down to two parts.
>
> The first is file-level and changes over time. A DV added later deletes
> some rows, and the accuracy of the original bounds are no longer
> guaranteed. The new file wide flag can help record the change.
>
> The second is column-level and static. Variable-length columns like String
> and Binary can have truncated bounds, so we cannot make them reliably
> tight, and we don't have a use case that needs their tightness today.
> That's a property of the column type in the schema, so it doesn't need
> additional metadata for recording.
>
> So moving the tightness boolean from per-column to per-entry (tracked_file
> in v4) simplifies the bookkeeping for writers without giving up anything we
> care about.
>
> One note on upgrading the existing tables to v4. Since we want that to
> stay a lightweight metadata operation with no manifest rewrite, tightness
> is simply undefined for existing entries. Today the merge-on-read table
> does not have tight bounds. Once position deletes and DVs are rewritten to
> be colocated with the data file in a single tracked-file entry, tightness
> becomes a well-defined property of that entry, and a writer at that point
> can set the flag accordingly.
>
> Thanks,
> Hongyue
>
> On Tue, Oct 6, 2026 at 12:23 AM Eduard Tudenhöfner <
> [email protected]> wrote:
>
>> The file-wide flag seems to make sense to me, so +1 for that
>>
>> On Mon, Oct 5, 2026 at 10:22 PM Daniel Weeks <[email protected]> wrote:
>>
>>> +1 to Russell's analysis here.
>>>
>>> While we can construct hypothetical use cases where there may be
>>> benefit, it seems like those use cases are unlikely and/or narrow.
>>>
>>> I'd also suggest going with the file-wide flag.
>>>
>>> -Dan
>>>
>>> On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer <
>>> [email protected]> wrote:
>>>
>>>> I've been thinking through the patterns where per-column tightness
>>>> would matter.
>>>>
>>>>    1.
>>>>
>>>>    *Selectively marking variable-length columns tight*. I don't have a
>>>>    use case for this yet. A file-level flag can still mean every 
>>>> non-truncated
>>>>    column is tight.
>>>>    2.
>>>>
>>>>    *Selectively re-tightening a few columns, for example after a
>>>>    column update*. Dan's point from the discussion applies here: a
>>>>    workload whose DVs invalidate stats will also invalidate new column 
>>>> stats.
>>>>    A table left in a state where only some stats are tight due to an update
>>>>    seems unlikely.
>>>>    3.
>>>>
>>>>    *A DV writer that can write tight stats*. This is the pattern that
>>>>    matters to me. The writer repairs bounds while it deletes, and it can
>>>>    repair every column it tracks. We still need an explicit flag, because
>>>>    otherwise readers treat any attached DV as wide and ignore the repaired
>>>>    bounds. One file-level flag covers this.
>>>>    4.
>>>>
>>>>    *A command that restores tight stats when the DV writer cannot.*
>>>>    This is worth supporting when deletes are infrequent: write the DVs,
>>>>    analyze once, and use the tight bounds until the next DV. If we 
>>>> frequently
>>>>    write DV's then we are in the same bad situation as #2. Either way the
>>>>    restore is all-or-nothing. I thought for a bit that there could be a
>>>>    command that only restores stats for select columns, this could make 
>>>> sense
>>>>    with column update files but also feels like unecessary complexity.
>>>>
>>>> Patterns 3 and 4 are the realistic ones, and both only need the
>>>> file-wide flag. So I'm kind of leaning in that direction now.
>>>>
>>>> On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]> wrote:
>>>>
>>>>> Hi, everyone -
>>>>>
>>>>> In v4, we introduced the notion of *tightness* in stats. Tightness
>>>>> indicates whether the upper/lower bounds truly exist in the live rows of a
>>>>> file. If the stats are tight, then engines can correctly answer many
>>>>> queries purely by consulting the metadata.
>>>>>
>>>>> However, when a DV is attached to a file, the stats are assumed to be
>>>>> non-tight, so these  metadata-only query optimizations won't work anymore.
>>>>>
>>>>> So we want a way to represent stats tightness even in the presence of
>>>>> a DV. I have a short writeup
>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0>
>>>>> of the problem and a few possible solutions. We discussed this today
>>>>> at the v4 adaptive metadata tree sync, and I wanted to continue the
>>>>> discussion.
>>>>>
>>>>> Please take a look at the doc
>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r>
>>>>> and let us know what you think by leaving a comment in the doc or in this
>>>>> email thread.
>>>>>
>>>>> Best,
>>>>> Anoop
>>>>>
>>>>

Reply via email to