The file-wide flag seems to make sense to me, so +1 for that On Mon, Oct 5, 2026 at 10:22 PM Daniel Weeks <[email protected]> wrote:
> +1 to Russell's analysis here. > > While we can construct hypothetical use cases where there may be benefit, > it seems like those use cases are unlikely and/or narrow. > > I'd also suggest going with the file-wide flag. > > -Dan > > On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer <[email protected]> > wrote: > >> I've been thinking through the patterns where per-column tightness would >> matter. >> >> 1. >> >> *Selectively marking variable-length columns tight*. I don't have a >> use case for this yet. A file-level flag can still mean every >> non-truncated >> column is tight. >> 2. >> >> *Selectively re-tightening a few columns, for example after a column >> update*. Dan's point from the discussion applies here: a workload >> whose DVs invalidate stats will also invalidate new column stats. A table >> left in a state where only some stats are tight due to an update seems >> unlikely. >> 3. >> >> *A DV writer that can write tight stats*. This is the pattern that >> matters to me. The writer repairs bounds while it deletes, and it can >> repair every column it tracks. We still need an explicit flag, because >> otherwise readers treat any attached DV as wide and ignore the repaired >> bounds. One file-level flag covers this. >> 4. >> >> *A command that restores tight stats when the DV writer cannot.* This >> is worth supporting when deletes are infrequent: write the DVs, analyze >> once, and use the tight bounds until the next DV. If we frequently write >> DV's then we are in the same bad situation as #2. Either way the restore >> is >> all-or-nothing. I thought for a bit that there could be a command that >> only >> restores stats for select columns, this could make sense with column >> update >> files but also feels like unecessary complexity. >> >> Patterns 3 and 4 are the realistic ones, and both only need the file-wide >> flag. So I'm kind of leaning in that direction now. >> >> On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]> wrote: >> >>> Hi, everyone - >>> >>> In v4, we introduced the notion of *tightness* in stats. Tightness >>> indicates whether the upper/lower bounds truly exist in the live rows of a >>> file. If the stats are tight, then engines can correctly answer many >>> queries purely by consulting the metadata. >>> >>> However, when a DV is attached to a file, the stats are assumed to be >>> non-tight, so these metadata-only query optimizations won't work anymore. >>> >>> So we want a way to represent stats tightness even in the presence of a >>> DV. I have a short writeup >>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0> >>> of the problem and a few possible solutions. We discussed this today at >>> the v4 adaptive metadata tree sync, and I wanted to continue the discussion. >>> >>> Please take a look at the doc >>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r> >>> and let us know what you think by leaving a comment in the doc or in this >>> email thread. >>> >>> Best, >>> Anoop >>> >>
