Hi Yunhong, +1 directionally. I read through both the FIP and #2981, and had a few questions and comments.
Bindings: arrow-rs's parquet-variant covers much of the spec mechanics, so the encode/decode/shred layer has a reasonable upstream path. What each binding would still need to port is the policy layer Fluss adds on top: inferrer, statistics collector, thresholds, candidate ranking, and tie-breaking. With shredding decisions made per writer (as the writer.variant.shredding.* configuration implies), each port has to produce compatible decisions with the Java inferrer for the same input, otherwise batches may fragment further. The doc in the PR itself acknowledges this can produce heterogeneous batches across Java writers, multi-language writers amplify that. A server-coordinated alternative would sidestep most of this. Is writer-side the v1 commitment, or Is server-coordination on the post-1.0 roadmap? Asking because the answer changes the binding strategy. Lance: the PR covers Iceberg and Paimon, but I did not see corresponding changes for fluss-lake-lance. Lance upstream Variant support also still appears to be in progress, and ArrowDataConverter.convertToNonShaded does not seem to handle StructVector. If a user tiers Variant data to Lance today, what happens? We might want to be explicit here. Arrow Java migration: the Java side implements custom shredding because Arrow Java does not yet provide native shredding support. When upstream catches up, do we intend to converge on Arrow Java's implementation, or keep the Fluss implementation as the canonical one for stability? Doc/code alignment: the design doc weighs writer-independent vs server-coordinated shredding as still open, but VariantShreddingManager and the writer.variant.shredding.* config commit to writer-independent. One observation while tracing the WAL path: walBuilder.append re-encodes Variant via ArrowShreddedVariantWriter, so UPDATE_BEFORE on an old-layout row gets re-shredded under the current policy. That keeps BEFORE/AFTER consistent, but buildAndWriteResidual appears to allocate per row: two ArrayLists, Arrays.copyOfRange per non-shredded field, and a fresh byte[] from encodeObject. That means UPDATE pays this twice, while INSERT pays it once. Not a blocker, but it is a non-obvious write-path cost from reading the FIP. -- Anton вт, 28 апр. 2026 г. в 02:29, yunhong Zheng <[email protected]>: > > Hi all, > > Semi-structured data (JSON, event payloads, metadata bags) is > ubiquitous in modern data pipelines. In the current Fluss > implementation, users are forced to store such data in STRING columns > and parse it at query time. This approach suffers from several > fundamental problems: no compact binary encoding, no field-level > access, and no ability to leverage Fluss's existing columnar > optimizations (column pruning, predicate pushdown). > > Meanwhile, the industry is converging on the Parquet Variant Binary > Encoding specification as the standard for semi-structured data — > Apache Spark 4.0, Flink 2.1 (FLIP-521), Paimon, and Parquet itself are > all adopting it. As a streaming storage layer bridging Flink and > Paimon in stream-lake architectures, Fluss needs native Variant > support to enable seamless data flow across the ecosystem. > > So, I'd like to propose FIP-36: Support Variant Type and Shredding[1]. > This proposal introduces a three-layer design: > > 1. Variant Binary Encoding: Adopt the Parquet Variant spec as the > canonical binary representation, enabling O(log n) field lookup and > zero-copy nested extraction. > > 2. Arrow Columnar Storage: Store Variant values in Arrow StructVector > with decomposed metadata/value children, integrating with Fluss's > existing zero-copy RecordBatch infrastructure. > > 3. Shredding for Query Optimization: Automatically extract frequently > accessed fields into typed sub-columns within the Variant > StructVector, enabling server-side sub-column pruning on > semi-structured data — all transparent to users. > > Any feedback and suggestions on this proposal are welcome! > > [1]: > https://cwiki.apache.org/confluence/display/FLUSS/FIP-36%3A+Support+Variant+Type+and+Shredding > > Regards, > Yunhong (Swuferhong)
