+1 to Russell's analysis here. While we can construct hypothetical use cases where there may be benefit, it seems like those use cases are unlikely and/or narrow.
I'd also suggest going with the file-wide flag. -Dan On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer <[email protected]> wrote: > I've been thinking through the patterns where per-column tightness would > matter. > > 1. > > *Selectively marking variable-length columns tight*. I don't have a > use case for this yet. A file-level flag can still mean every non-truncated > column is tight. > 2. > > *Selectively re-tightening a few columns, for example after a column > update*. Dan's point from the discussion applies here: a workload > whose DVs invalidate stats will also invalidate new column stats. A table > left in a state where only some stats are tight due to an update seems > unlikely. > 3. > > *A DV writer that can write tight stats*. This is the pattern that > matters to me. The writer repairs bounds while it deletes, and it can > repair every column it tracks. We still need an explicit flag, because > otherwise readers treat any attached DV as wide and ignore the repaired > bounds. One file-level flag covers this. > 4. > > *A command that restores tight stats when the DV writer cannot.* This > is worth supporting when deletes are infrequent: write the DVs, analyze > once, and use the tight bounds until the next DV. If we frequently write > DV's then we are in the same bad situation as #2. Either way the restore is > all-or-nothing. I thought for a bit that there could be a command that only > restores stats for select columns, this could make sense with column update > files but also feels like unecessary complexity. > > Patterns 3 and 4 are the realistic ones, and both only need the file-wide > flag. So I'm kind of leaning in that direction now. > > On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]> wrote: > >> Hi, everyone - >> >> In v4, we introduced the notion of *tightness* in stats. Tightness >> indicates whether the upper/lower bounds truly exist in the live rows of a >> file. If the stats are tight, then engines can correctly answer many >> queries purely by consulting the metadata. >> >> However, when a DV is attached to a file, the stats are assumed to be >> non-tight, so these metadata-only query optimizations won't work anymore. >> >> So we want a way to represent stats tightness even in the presence of a >> DV. I have a short writeup >> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0> >> of the problem and a few possible solutions. We discussed this today at >> the v4 adaptive metadata tree sync, and I wanted to continue the discussion. >> >> Please take a look at the doc >> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r> >> and let us know what you think by leaving a comment in the doc or in this >> email thread. >> >> Best, >> Anoop >> >
