+1 on the file-wide flag. >From what I can tell, whether a given lower/upper bound is tight comes down to two parts.
The first is file-level and changes over time. A DV added later deletes some rows, and the accuracy of the original bounds are no longer guaranteed. The new file wide flag can help record the change. The second is column-level and static. Variable-length columns like String and Binary can have truncated bounds, so we cannot make them reliably tight, and we don't have a use case that needs their tightness today. That's a property of the column type in the schema, so it doesn't need additional metadata for recording. So moving the tightness boolean from per-column to per-entry (tracked_file in v4) simplifies the bookkeeping for writers without giving up anything we care about. One note on upgrading the existing tables to v4. Since we want that to stay a lightweight metadata operation with no manifest rewrite, tightness is simply undefined for existing entries. Today the merge-on-read table does not have tight bounds. Once position deletes and DVs are rewritten to be colocated with the data file in a single tracked-file entry, tightness becomes a well-defined property of that entry, and a writer at that point can set the flag accordingly. Thanks, Hongyue On Tue, Oct 6, 2026 at 12:23 AM Eduard Tudenhöfner <[email protected]> wrote: > The file-wide flag seems to make sense to me, so +1 for that > > On Mon, Oct 5, 2026 at 10:22 PM Daniel Weeks <[email protected]> wrote: > >> +1 to Russell's analysis here. >> >> While we can construct hypothetical use cases where there may be benefit, >> it seems like those use cases are unlikely and/or narrow. >> >> I'd also suggest going with the file-wide flag. >> >> -Dan >> >> On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer < >> [email protected]> wrote: >> >>> I've been thinking through the patterns where per-column tightness would >>> matter. >>> >>> 1. >>> >>> *Selectively marking variable-length columns tight*. I don't have a >>> use case for this yet. A file-level flag can still mean every >>> non-truncated >>> column is tight. >>> 2. >>> >>> *Selectively re-tightening a few columns, for example after a column >>> update*. Dan's point from the discussion applies here: a workload >>> whose DVs invalidate stats will also invalidate new column stats. A table >>> left in a state where only some stats are tight due to an update seems >>> unlikely. >>> 3. >>> >>> *A DV writer that can write tight stats*. This is the pattern that >>> matters to me. The writer repairs bounds while it deletes, and it can >>> repair every column it tracks. We still need an explicit flag, because >>> otherwise readers treat any attached DV as wide and ignore the repaired >>> bounds. One file-level flag covers this. >>> 4. >>> >>> *A command that restores tight stats when the DV writer cannot.* >>> This is worth supporting when deletes are infrequent: write the DVs, >>> analyze once, and use the tight bounds until the next DV. If we >>> frequently >>> write DV's then we are in the same bad situation as #2. Either way the >>> restore is all-or-nothing. I thought for a bit that there could be a >>> command that only restores stats for select columns, this could make >>> sense >>> with column update files but also feels like unecessary complexity. >>> >>> Patterns 3 and 4 are the realistic ones, and both only need the >>> file-wide flag. So I'm kind of leaning in that direction now. >>> >>> On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]> wrote: >>> >>>> Hi, everyone - >>>> >>>> In v4, we introduced the notion of *tightness* in stats. Tightness >>>> indicates whether the upper/lower bounds truly exist in the live rows of a >>>> file. If the stats are tight, then engines can correctly answer many >>>> queries purely by consulting the metadata. >>>> >>>> However, when a DV is attached to a file, the stats are assumed to be >>>> non-tight, so these metadata-only query optimizations won't work anymore. >>>> >>>> So we want a way to represent stats tightness even in the presence of a >>>> DV. I have a short writeup >>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0> >>>> of the problem and a few possible solutions. We discussed this today >>>> at the v4 adaptive metadata tree sync, and I wanted to continue the >>>> discussion. >>>> >>>> Please take a look at the doc >>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r> >>>> and let us know what you think by leaving a comment in the doc or in this >>>> email thread. >>>> >>>> Best, >>>> Anoop >>>> >>>
