JingsongLi commented on code in PR #858:
URL: https://github.com/apache/paimon-rust/pull/858#discussion_r4068000247
##########
crates/paimon/src/spec/avro/manifest_entry_decode.rs:
##########
@@ -162,11 +162,183 @@ where
normalize_partition(partition),
bucket.unwrap_or(0),
total_buckets.unwrap_or(0),
- file.unwrap_or_else(default_data_file_meta),
+ file.ok_or_else(missing_file_metadata)?,
version.unwrap_or(0),
)))
}
+/// Borrowed view of the manifest-entry fields needed to count rows per
partition.
+/// Everything else (key/value stats, min/max keys, ...) is skipped in place,
so
+/// decoding never materializes the per-file statistics that dominate manifest
size.
+/// Ordinary files allocate nothing; files with `extra_files` allocate only
that list.
+#[derive(Debug)]
+pub(crate) struct SlimManifestEntry<'a> {
+ pub kind: FileKind,
+ pub partition: &'a [u8],
+ pub bucket: i32,
+ pub level: i32,
+ pub file_name: &'a str,
+ pub row_count: i64,
+ pub first_row_id: Option<i64>,
+ pub extra_files: Vec<&'a str>,
+ pub embedded_index: Option<&'a [u8]>,
+ pub external_path: Option<&'a str>,
+}
+
+/// Decode one manifest entry as a [`SlimManifestEntry`] borrowing from the
block.
+pub(crate) fn decode_slim_manifest_entry<'a>(
Review Comment:
Non-blocking design suggestion: please keep the borrowed slim
representation, but consider sharing the record walk rather than maintaining a
third field decoder. `ManifestEntry::decode`,
`decode_manifest_entries_filtered`, and this function each interpret the
writer-schema fields and their union/null/default behavior independently, so a
future manifest schema change could update only one path. A shared field
visitor with separate full and count collectors would retain zero-copy skipping
of column statistics. The index-manifest path already shares its trickiest
DV-array parsing through `visit_nullable_dv_ranges`; the two `visit_slim_*` OCF
block loops could also use one helper. Differential tests for reordered,
nullable, and unknown fields would protect the shared path.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]