MisterRaindrop opened a new issue, #1989: URL: https://github.com/apache/cloudberry/issues/1989
### Summary Since #1951 the Parquet reader in `contrib/datalake_fdw` matches a table's columns to a file's by Iceberg field id (`PARQUET:field_id` in the Parquet schema). A file whose columns carry no field id can never be matched: every projected column reads as NULL. That is correct per spec for a file the table never wrote, and wrong for the one case the spec covers: tables created over existing Parquet data (`add_files`, migrated Hive tables), whose files predate the ids. ### What the spec says Iceberg resolves such files through the table property `schema.name-mapping.default`: a JSON mapping from field ids to the column names to look for in the file, including nested names and multiple names per id (for renamed columns). A reader applies it only to columns that have no field id. ### What has to happen - The metadata engine has to surface the property to the access method. - `ProjectionSet` (`format/format.h`) needs a way to hand the reader a name mapping alongside the field ids -- a second array of names per id, or a pointer to a parsed mapping. - `parquet_read.cpp`'s `parquet_project()` matches by id first and falls back to the mapping for id-less columns. Nothing changes in the decoder. - A file with neither ids nor a mapping match stays NULL, as now. Deferred from #1951 on purpose: the framework there is settled, and the property needs the metadata engine, which is a later PR. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
