Hi Jörn, Thanks for the clarification. I'm not against introducing an UNKNOWN column order (that would mean that any min/max values are not usable for the related columns). However, I still feel that writing columns/statistics etc. without the proper metadata (logical type, column order etc.) is not a good approach.
Cheers, Gabor Jörn Horstmann <[email protected]> ezt írta (időpont: 2026. szept. 14., H, 11:48): > Thanks Gabor, > > The difference I wanted to highlight is that in this usecase, new > logical types can have a local impact on individual columns, but new > column orders would have an impact on all the columns in the file. There > is no way to mark only a single column as having an unknown or undefined > column order. > > The arrow-rs implementation for example handles this by having an > internal "unknown" variant, which it is then not able to serialize > again. But I would argue that, if it is fine for a reader to ignore > unknown column orders or logical types, then it should also be fine if > they are ignored after such a preprocessing or indirect reading. > > Cheers, > Jörn > > On 9/10/26 09:39, Gábor Szádovszky wrote: > > Hi Jörn, > > > > Your mail landed in my spam folder, so let me write an answer to bring up > > this topic for the others who might not see it. > > > > I don't think this is only a problem with column orders. If your rewriter > > do not know the related logical type, additional statistics etc. in the > > footer, you are losing crucial information even if you can bypass the raw > > data with the related encodings. Let's say your file has a shredded > VARIANT > > column while your rewriter do not know the VARIANT logical type. In this > > case, in the rewritten file, it would be a huge column structure of > > structs, lists with different primitives in them. Without the VARIANT > > logical type (and the version number in it) this whole bunch of data is > > useless. > > > > I think, a rewriter shall at least "understand" the related footer so it > > can read/write it properly. Should fail otherwise. > > > > Cheers, > > Gabor > > > > Jörn Horstmann <[email protected]> ezt írta (időpont: 2026. szept. > 9., > > Sze, 10:29): > > > >> Hello everyone, > >> > >> Let's assume I want to create a program that rewrites and existing > >> parquet file, for example by adding key/value metadata or adding > >> additional columns. The input file might use some new features, like > >> encodings, logical types or column orders. > >> > >> Encodings are simple thrift enums, in most implementations these are > >> handled as a simple wrapper around their id and not as static enums. So > >> as long as the corresponding column data is copied byte-identical, > >> roundtripping should work. > >> > >> Logical types are optional in the thrift structures, so new variants > >> should be skipped by the reading side, and written as absent values. > >> Consumers of a file might then have to cast the values, but could > >> otherwise still process the file. > >> > >> Column orders however are contained in a list. Readers commonly map > >> unknown values to an internal "Unknown" variant which it is then not > >> able to serialize to thrift again. When rewriting a file that contained > >> an unknown column order, we would either have to skip the whole list, or > >> write a potentially incorrect order for some columns. Missing column > >> orders then probably make statistics in that file unusable. > >> > >> I wonder how any existing parquet rewriters handle these problems. And > >> Maybe it would make sense to define a canonical "UNKNOWN" ColumnOrder > >> that could be written in such cases. > >> > >> Cheers, > >> Jörn > >> > >> >
