Yes, the question is whether we want to persist _pos into the column files.
We earlier concluded that even though we went with the dense representation
we persist _pos. However, I think re-opening the question is reasonable,
because that's just extra noise ATM, and we shouldn't write that field into
column files. Reading the _pos column would still work regardless if we
include the base file or not to the query.

Best Regards,
Gabor


Russell Spitzer <[email protected]> ezt írta (időpont: 2026. aug.
26., Sze, 16:21):

> What is the actual argument here? I think having a persisted field doesn't
> make sense since we expect perfect alignment. We would expect reading the
> file in isolation with the metadata _pos column should still work right?
>
> If we are just discussing removing a persisted value, I'm in favor of that.
>
> On Wed, Aug 26, 2026 at 8:37 AM Gábor Kaszab <[email protected]>
> wrote:
>
>> I hear you, and I share the same opinion. If we don't need such a field
>> then it's just extra unnecessary complexity to write it. I'm not entirely
>> convinced on the debugging use of the _pos field. Would be beneficial to
>> reduce unnecessary noise and confusion by not adding the _pos field.
>>
>> Let's discuss this on the next sync! In the meantime, opinions are
>> welcome here too.
>>
>> Thanks,
>> Gabor
>>
>> Andrei Tserakhau via dev <[email protected]> ezt írta (időpont:
>> 2026. aug. 26., Sze, 15:32):
>>
>>> +1 on this question.
>>>
>>> Right now `_pos` column feels more like debug leftovers, it bring some
>>> confusion for read-side weather it's expected to be readed or not.
>>>
>>> I think removing it would make implementation easier.
>>>
>>> Best,
>>> Andrei
>>>
>>> On Wed, Aug 26, 2026 at 2:42 PM Leonid Lygin via dev <
>>> [email protected]> wrote:
>>>
>>>> Thanks for the quick response!
>>>>
>>>> My biggest concern with `_pos` is not performance but rather clarity
>>>> and implementation divergence:
>>>>
>>>> 1. including `_pos` is redundant, and (at least for me) provokes a
>>>> re-read of the row alignment section — "why include `_pos` if files
>>>> are fully aligned?";
>>>> 2. having `_pos` fully duplicate the row position, there are two
>>>> different legal ways to implement reads: either positionally, or using
>>>> `_pos`.
>>>>
>>>> On Wed, Aug 26, 2026 at 2:35 PM Gábor Kaszab <[email protected]>
>>>> wrote:
>>>> >
>>>> > Hey All,
>>>> >
>>>> > Thanks for bringing this up! (for me the initial mail went to spam,
>>>> though...)
>>>> >
>>>> > Technically, with the dense representation we don't really need the
>>>> _pos column in the column files, unless for troubleshooting. While checking
>>>> the row counts is good, if they don't match we might get a better
>>>> understanding on what the writer missed writing if we had the _pos col,
>>>> also the order could be verified.
>>>> >
>>>> > Apart from debugging, I think either way is just fine. An additional
>>>> detail to consider is that according to my experiments, there isn't really
>>>> any storage cost for writing the _pos with delta encoding (e.g. with
>>>> Parquet V2). So the conclusion was that since it comes for free, and might
>>>> help for debugging, why not write it.
>>>> >
>>>> > Should we reopen this question? Any further feedback is welcome.
>>>> >
>>>> > Best Regards,
>>>> > Gabor
>>>> >
>>>> > Leonid Lygin via dev <[email protected]> ezt írta (időpont:
>>>> 2026. aug. 26., Sze, 14:10):
>>>> >>
>>>> >> Definitely agree that including `_pos` raises questions.
>>>> >>
>>>> >> If "debugging" is to be understood as figuring out if the column
>>>> files
>>>> >> have gaps -- just checking the row counts is good enough for that. Is
>>>> >> there a lot to be gained from figuring out where exactly the gap is
>>>> >> occurring?
>>>> >>
>>>> >> On Mon, Aug 24, 2026 at 1:57 PM Marco Kroll
>>>> >> <[email protected]> wrote:
>>>> >> >
>>>> >> > Hi all,
>>>> >> >
>>>> >> > I just saw the agenda [1] for tomorrow's (2026-08-25) sync and
>>>> want to +1 the `_pos` column topic.
>>>> >> > My understanding is that this column exists for two reasons:
>>>> >> > 1. debugging
>>>> >> > 2. detect if writers skipped deleted rows
>>>> >> >
>>>> >> > My take is that using the dense Null filled representation
>>>> addresses both of these issues.
>>>> >> > It implicitly encodes the position, very much like for deletion
>>>> vectors and since all rows need to be present, comparing the row count of
>>>> the base file with the column file can be used to verify that all rows were
>>>> written.
>>>> >> >
>>>> >> > The main thing to add to the doc would be that the row order must
>>>> be identical to the base file.
>>>> >> >
>>>> >> > Best
>>>> >> > Marco
>>>> >> >
>>>> >> > [1]:
>>>> https://docs.google.com/document/d/1Bd7JVzgajA8-DozzeEE24mID_GLuz6iwj0g4TlcVJcs/edit?tab=t.jvm7iiiulf8q#heading=h.rbisiun18esp
>>>>
>>>

Reply via email to