Le 23/09/2026 à 16:26, Rok Mihevc a écrit :
Well, Parquet is not a single-purpose format.

Vector stores can choose whatever policy fits their usage, but we're
talking about a general-purpose file format that aims to be broadly
applicable.

What is the gate for Parquet to define something as a logical type?

There is no "gate". But we should act in the broader community's interest. A limitation that is only motivated by the needs of a single segment of the user base goes against that goal.

Why not? Why shouldn't it be used for storing, for example, NumPy tensor
data (which can contain NaNs and infinites)?

A multidimensional arrays logical type might come with other
requirements (shapes, strides, dimension names), that vector
applications might not have.

Well, first, you can use NumPy just for 1D data, and I'm sure some people do that. Second, strides wouldn't be saved in Parquet where they are meaningless. Third, all of this can be conveyed as additional key-value metadata.

It does not make sense to have *both* a "specialist" Vector type and a
"general-purpose" FixedSizeList type. Parquet types are a
general-purpose vocabulary from which you can build up more specialized
applications. Specialist types can be left to domain-specific SW users
of Parquet.

(for example, Arrow implementations serialize the Arrow schema in
Parquet metadata so that the Arrow schema can be rebuilt at read time,
without mandating that every Arrow datatype has a Parquet equivalent)

That does seem to work nicely for Arrow [0]. Should we revisit the Parquet
extensions discussion [1][2]?

I think we should, and then specialized communities such as ML can define their own specialized extension types (with arbitrary semantics) on top of regular Parquet logical/physical types.

Regards

Antoine.


Reply via email to