Yes, we are discussing removing the persisted _pos column. I believe it was a remnant from the older design where we didn't have row alignment between column files. I think it's no longer necessary and we should just remove it.
> We would expect reading the file in isolation with the metadata _pos column should still work right? Yes, of course. That should continue to work. It is a pre-req for inherited row IDs anyway. Best, Anoop On Wed, Aug 26, 2026 at 7:21 AM Russell Spitzer <[email protected]> wrote: > What is the actual argument here? I think having a persisted field doesn't > make sense since we expect perfect alignment. We would expect reading the > file in isolation with the metadata _pos column should still work right? > > If we are just discussing removing a persisted value, I'm in favor of that. > > On Wed, Aug 26, 2026 at 8:37 AM Gábor Kaszab <[email protected]> > wrote: > >> I hear you, and I share the same opinion. If we don't need such a field >> then it's just extra unnecessary complexity to write it. I'm not entirely >> convinced on the debugging use of the _pos field. Would be beneficial to >> reduce unnecessary noise and confusion by not adding the _pos field. >> >> Let's discuss this on the next sync! In the meantime, opinions are >> welcome here too. >> >> Thanks, >> Gabor >> >> Andrei Tserakhau via dev <[email protected]> ezt írta (időpont: >> 2026. aug. 26., Sze, 15:32): >> >>> +1 on this question. >>> >>> Right now `_pos` column feels more like debug leftovers, it bring some >>> confusion for read-side weather it's expected to be readed or not. >>> >>> I think removing it would make implementation easier. >>> >>> Best, >>> Andrei >>> >>> On Wed, Aug 26, 2026 at 2:42 PM Leonid Lygin via dev < >>> [email protected]> wrote: >>> >>>> Thanks for the quick response! >>>> >>>> My biggest concern with `_pos` is not performance but rather clarity >>>> and implementation divergence: >>>> >>>> 1. including `_pos` is redundant, and (at least for me) provokes a >>>> re-read of the row alignment section — "why include `_pos` if files >>>> are fully aligned?"; >>>> 2. having `_pos` fully duplicate the row position, there are two >>>> different legal ways to implement reads: either positionally, or using >>>> `_pos`. >>>> >>>> On Wed, Aug 26, 2026 at 2:35 PM Gábor Kaszab <[email protected]> >>>> wrote: >>>> > >>>> > Hey All, >>>> > >>>> > Thanks for bringing this up! (for me the initial mail went to spam, >>>> though...) >>>> > >>>> > Technically, with the dense representation we don't really need the >>>> _pos column in the column files, unless for troubleshooting. While checking >>>> the row counts is good, if they don't match we might get a better >>>> understanding on what the writer missed writing if we had the _pos col, >>>> also the order could be verified. >>>> > >>>> > Apart from debugging, I think either way is just fine. An additional >>>> detail to consider is that according to my experiments, there isn't really >>>> any storage cost for writing the _pos with delta encoding (e.g. with >>>> Parquet V2). So the conclusion was that since it comes for free, and might >>>> help for debugging, why not write it. >>>> > >>>> > Should we reopen this question? Any further feedback is welcome. >>>> > >>>> > Best Regards, >>>> > Gabor >>>> > >>>> > Leonid Lygin via dev <[email protected]> ezt írta (időpont: >>>> 2026. aug. 26., Sze, 14:10): >>>> >> >>>> >> Definitely agree that including `_pos` raises questions. >>>> >> >>>> >> If "debugging" is to be understood as figuring out if the column >>>> files >>>> >> have gaps -- just checking the row counts is good enough for that. Is >>>> >> there a lot to be gained from figuring out where exactly the gap is >>>> >> occurring? >>>> >> >>>> >> On Mon, Aug 24, 2026 at 1:57 PM Marco Kroll >>>> >> <[email protected]> wrote: >>>> >> > >>>> >> > Hi all, >>>> >> > >>>> >> > I just saw the agenda [1] for tomorrow's (2026-08-25) sync and >>>> want to +1 the `_pos` column topic. >>>> >> > My understanding is that this column exists for two reasons: >>>> >> > 1. debugging >>>> >> > 2. detect if writers skipped deleted rows >>>> >> > >>>> >> > My take is that using the dense Null filled representation >>>> addresses both of these issues. >>>> >> > It implicitly encodes the position, very much like for deletion >>>> vectors and since all rows need to be present, comparing the row count of >>>> the base file with the column file can be used to verify that all rows were >>>> written. >>>> >> > >>>> >> > The main thing to add to the doc would be that the row order must >>>> be identical to the base file. >>>> >> > >>>> >> > Best >>>> >> > Marco >>>> >> > >>>> >> > [1]: >>>> https://docs.google.com/document/d/1Bd7JVzgajA8-DozzeEE24mID_GLuz6iwj0g4TlcVJcs/edit?tab=t.jvm7iiiulf8q#heading=h.rbisiun18esp >>>> >>>
