The addition of `_pos` column was not a remnant of previous design, but a
conscious choice we made during one of the column updates sync. I'm in
favor of removing it if it makes implementation easier.

~ Anurag

On Wed, Aug 26, 2026 at 7:31 AM Gábor Kaszab <[email protected]> wrote:

> Yes, the question is whether we want to persist _pos into the column
> files. We earlier concluded that even though we went with the dense
> representation we persist _pos. However, I think re-opening the question is
> reasonable, because that's just extra noise ATM, and we shouldn't write
> that field into column files. Reading the _pos column would still work
> regardless if we include the base file or not to the query.
>
> Best Regards,
> Gabor
>
>
> Russell Spitzer <[email protected]> ezt írta (időpont: 2026. aug.
> 26., Sze, 16:21):
>
>> What is the actual argument here? I think having a persisted field
>> doesn't make sense since we expect perfect alignment. We would expect
>> reading the file in isolation with the metadata _pos column should still
>> work right?
>>
>> If we are just discussing removing a persisted value, I'm in favor of
>> that.
>>
>> On Wed, Aug 26, 2026 at 8:37 AM Gábor Kaszab <[email protected]>
>> wrote:
>>
>>> I hear you, and I share the same opinion. If we don't need such a field
>>> then it's just extra unnecessary complexity to write it. I'm not entirely
>>> convinced on the debugging use of the _pos field. Would be beneficial to
>>> reduce unnecessary noise and confusion by not adding the _pos field.
>>>
>>> Let's discuss this on the next sync! In the meantime, opinions are
>>> welcome here too.
>>>
>>> Thanks,
>>> Gabor
>>>
>>> Andrei Tserakhau via dev <[email protected]> ezt írta (időpont:
>>> 2026. aug. 26., Sze, 15:32):
>>>
>>>> +1 on this question.
>>>>
>>>> Right now `_pos` column feels more like debug leftovers, it bring some
>>>> confusion for read-side weather it's expected to be readed or not.
>>>>
>>>> I think removing it would make implementation easier.
>>>>
>>>> Best,
>>>> Andrei
>>>>
>>>> On Wed, Aug 26, 2026 at 2:42 PM Leonid Lygin via dev <
>>>> [email protected]> wrote:
>>>>
>>>>> Thanks for the quick response!
>>>>>
>>>>> My biggest concern with `_pos` is not performance but rather clarity
>>>>> and implementation divergence:
>>>>>
>>>>> 1. including `_pos` is redundant, and (at least for me) provokes a
>>>>> re-read of the row alignment section — "why include `_pos` if files
>>>>> are fully aligned?";
>>>>> 2. having `_pos` fully duplicate the row position, there are two
>>>>> different legal ways to implement reads: either positionally, or using
>>>>> `_pos`.
>>>>>
>>>>> On Wed, Aug 26, 2026 at 2:35 PM Gábor Kaszab <[email protected]>
>>>>> wrote:
>>>>> >
>>>>> > Hey All,
>>>>> >
>>>>> > Thanks for bringing this up! (for me the initial mail went to spam,
>>>>> though...)
>>>>> >
>>>>> > Technically, with the dense representation we don't really need the
>>>>> _pos column in the column files, unless for troubleshooting. While 
>>>>> checking
>>>>> the row counts is good, if they don't match we might get a better
>>>>> understanding on what the writer missed writing if we had the _pos col,
>>>>> also the order could be verified.
>>>>> >
>>>>> > Apart from debugging, I think either way is just fine. An additional
>>>>> detail to consider is that according to my experiments, there isn't really
>>>>> any storage cost for writing the _pos with delta encoding (e.g. with
>>>>> Parquet V2). So the conclusion was that since it comes for free, and might
>>>>> help for debugging, why not write it.
>>>>> >
>>>>> > Should we reopen this question? Any further feedback is welcome.
>>>>> >
>>>>> > Best Regards,
>>>>> > Gabor
>>>>> >
>>>>> > Leonid Lygin via dev <[email protected]> ezt írta (időpont:
>>>>> 2026. aug. 26., Sze, 14:10):
>>>>> >>
>>>>> >> Definitely agree that including `_pos` raises questions.
>>>>> >>
>>>>> >> If "debugging" is to be understood as figuring out if the column
>>>>> files
>>>>> >> have gaps -- just checking the row counts is good enough for that.
>>>>> Is
>>>>> >> there a lot to be gained from figuring out where exactly the gap is
>>>>> >> occurring?
>>>>> >>
>>>>> >> On Mon, Aug 24, 2026 at 1:57 PM Marco Kroll
>>>>> >> <[email protected]> wrote:
>>>>> >> >
>>>>> >> > Hi all,
>>>>> >> >
>>>>> >> > I just saw the agenda [1] for tomorrow's (2026-08-25) sync and
>>>>> want to +1 the `_pos` column topic.
>>>>> >> > My understanding is that this column exists for two reasons:
>>>>> >> > 1. debugging
>>>>> >> > 2. detect if writers skipped deleted rows
>>>>> >> >
>>>>> >> > My take is that using the dense Null filled representation
>>>>> addresses both of these issues.
>>>>> >> > It implicitly encodes the position, very much like for deletion
>>>>> vectors and since all rows need to be present, comparing the row count of
>>>>> the base file with the column file can be used to verify that all rows 
>>>>> were
>>>>> written.
>>>>> >> >
>>>>> >> > The main thing to add to the doc would be that the row order must
>>>>> be identical to the base file.
>>>>> >> >
>>>>> >> > Best
>>>>> >> > Marco
>>>>> >> >
>>>>> >> > [1]:
>>>>> https://docs.google.com/document/d/1Bd7JVzgajA8-DozzeEE24mID_GLuz6iwj0g4TlcVJcs/edit?tab=t.jvm7iiiulf8q#heading=h.rbisiun18esp
>>>>>
>>>>

Reply via email to