The addition of `_pos` column was not a remnant of previous design, but a conscious choice we made during one of the column updates sync. I'm in favor of removing it if it makes implementation easier.
~ Anurag On Wed, Aug 26, 2026 at 7:31 AM Gábor Kaszab <[email protected]> wrote: > Yes, the question is whether we want to persist _pos into the column > files. We earlier concluded that even though we went with the dense > representation we persist _pos. However, I think re-opening the question is > reasonable, because that's just extra noise ATM, and we shouldn't write > that field into column files. Reading the _pos column would still work > regardless if we include the base file or not to the query. > > Best Regards, > Gabor > > > Russell Spitzer <[email protected]> ezt írta (időpont: 2026. aug. > 26., Sze, 16:21): > >> What is the actual argument here? I think having a persisted field >> doesn't make sense since we expect perfect alignment. We would expect >> reading the file in isolation with the metadata _pos column should still >> work right? >> >> If we are just discussing removing a persisted value, I'm in favor of >> that. >> >> On Wed, Aug 26, 2026 at 8:37 AM Gábor Kaszab <[email protected]> >> wrote: >> >>> I hear you, and I share the same opinion. If we don't need such a field >>> then it's just extra unnecessary complexity to write it. I'm not entirely >>> convinced on the debugging use of the _pos field. Would be beneficial to >>> reduce unnecessary noise and confusion by not adding the _pos field. >>> >>> Let's discuss this on the next sync! In the meantime, opinions are >>> welcome here too. >>> >>> Thanks, >>> Gabor >>> >>> Andrei Tserakhau via dev <[email protected]> ezt írta (időpont: >>> 2026. aug. 26., Sze, 15:32): >>> >>>> +1 on this question. >>>> >>>> Right now `_pos` column feels more like debug leftovers, it bring some >>>> confusion for read-side weather it's expected to be readed or not. >>>> >>>> I think removing it would make implementation easier. >>>> >>>> Best, >>>> Andrei >>>> >>>> On Wed, Aug 26, 2026 at 2:42 PM Leonid Lygin via dev < >>>> [email protected]> wrote: >>>> >>>>> Thanks for the quick response! >>>>> >>>>> My biggest concern with `_pos` is not performance but rather clarity >>>>> and implementation divergence: >>>>> >>>>> 1. including `_pos` is redundant, and (at least for me) provokes a >>>>> re-read of the row alignment section — "why include `_pos` if files >>>>> are fully aligned?"; >>>>> 2. having `_pos` fully duplicate the row position, there are two >>>>> different legal ways to implement reads: either positionally, or using >>>>> `_pos`. >>>>> >>>>> On Wed, Aug 26, 2026 at 2:35 PM Gábor Kaszab <[email protected]> >>>>> wrote: >>>>> > >>>>> > Hey All, >>>>> > >>>>> > Thanks for bringing this up! (for me the initial mail went to spam, >>>>> though...) >>>>> > >>>>> > Technically, with the dense representation we don't really need the >>>>> _pos column in the column files, unless for troubleshooting. While >>>>> checking >>>>> the row counts is good, if they don't match we might get a better >>>>> understanding on what the writer missed writing if we had the _pos col, >>>>> also the order could be verified. >>>>> > >>>>> > Apart from debugging, I think either way is just fine. An additional >>>>> detail to consider is that according to my experiments, there isn't really >>>>> any storage cost for writing the _pos with delta encoding (e.g. with >>>>> Parquet V2). So the conclusion was that since it comes for free, and might >>>>> help for debugging, why not write it. >>>>> > >>>>> > Should we reopen this question? Any further feedback is welcome. >>>>> > >>>>> > Best Regards, >>>>> > Gabor >>>>> > >>>>> > Leonid Lygin via dev <[email protected]> ezt írta (időpont: >>>>> 2026. aug. 26., Sze, 14:10): >>>>> >> >>>>> >> Definitely agree that including `_pos` raises questions. >>>>> >> >>>>> >> If "debugging" is to be understood as figuring out if the column >>>>> files >>>>> >> have gaps -- just checking the row counts is good enough for that. >>>>> Is >>>>> >> there a lot to be gained from figuring out where exactly the gap is >>>>> >> occurring? >>>>> >> >>>>> >> On Mon, Aug 24, 2026 at 1:57 PM Marco Kroll >>>>> >> <[email protected]> wrote: >>>>> >> > >>>>> >> > Hi all, >>>>> >> > >>>>> >> > I just saw the agenda [1] for tomorrow's (2026-08-25) sync and >>>>> want to +1 the `_pos` column topic. >>>>> >> > My understanding is that this column exists for two reasons: >>>>> >> > 1. debugging >>>>> >> > 2. detect if writers skipped deleted rows >>>>> >> > >>>>> >> > My take is that using the dense Null filled representation >>>>> addresses both of these issues. >>>>> >> > It implicitly encodes the position, very much like for deletion >>>>> vectors and since all rows need to be present, comparing the row count of >>>>> the base file with the column file can be used to verify that all rows >>>>> were >>>>> written. >>>>> >> > >>>>> >> > The main thing to add to the doc would be that the row order must >>>>> be identical to the base file. >>>>> >> > >>>>> >> > Best >>>>> >> > Marco >>>>> >> > >>>>> >> > [1]: >>>>> https://docs.google.com/document/d/1Bd7JVzgajA8-DozzeEE24mID_GLuz6iwj0g4TlcVJcs/edit?tab=t.jvm7iiiulf8q#heading=h.rbisiun18esp >>>>> >>>>
