leaves12138 commented on code in PR #10047:
URL: https://github.com/apache/paimon/pull/10047#discussion_r4060116691


##########
paimon-python/pypaimon/read/table_read.py:
##########
@@ -923,6 +924,29 @@ def _deferred_blob_limit_may_prune(self, splits: 
List[Split]) -> bool:
                      or self._native_inline_blob_fields())
                 and not self._limit_covers_all_splits(splits))
 
+    def _native_pruning_blob_limit_supported(self) -> bool:
+        """Rust caps DE batches before payload resolution when no post-filter 
is needed.
+
+        A predicate on a managed BLOB or BLOB view can require payload I/O
+        before the output quota is known. Inline descriptors can still use
+        native reads if the predicate only references ordinary columns.
+        """
+        if not self.table.options.data_evolution_enabled():
+            return False
+        read_names = {field.name for field in self._scan_read_type}
+        if self.table.options.blob_view_fields() & read_names:
+            return False
+        if self.predicate is not None:
+            # Managed BLOBs are decoded by the physical file reader, before a
+            # residual filter. Inline descriptors are resolved later, so a
+            # predicate on ordinary columns can safely select rows first.
+            if self._deferred_blob_fields:
+                return False
+            from pypaimon.read.push_down_utils import predicate_field_names
+            if predicate_field_names(self.predicate) & 
self._native_inline_blob_fields():
+                return False
+        return True

Review Comment:
   [P2] Keep the managed-BLOB pruning-LIMIT fallback until the native 
remaining-quota handling is fixed.
   
   With no predicate, this returns `True` even when `_deferred_blob_fields` is 
nonempty, but apache/paimon-rust#896 caps physical batches using the initial 
LIMIT rather than the remaining quota. I verified two end-to-end failures 
against both PR heads: `LIMIT 3` with `read.batch-size=2` over five rows, and 
`LIMIT 3` over two successive appends containing two and three rows with the 
default batch size. Both native reads decode row 4's managed BLOB before 
discarding it. Corrupting only that unselected payload causes `Blob entry CRC32 
mismatch`, while `read.native.enabled=false` returns the expected first three 
rows.
   
   This removes the previous deferred-reader protection for these cases. Fixing 
the companion Rust implementation would make the gate safe; alternatively, keep 
managed-BLOB pruning LIMITs on the existing Python path for now and enable only 
the safe inline-descriptor case. Please add cross-batch and cross-split 
payload-I/O regressions, not just assertions on the returned row count.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to