JingsongLi commented on code in PR #9954:
URL: https://github.com/apache/paimon/pull/9954#discussion_r4046642240


##########
paimon-python/pypaimon/read/table_read.py:
##########
@@ -400,24 +599,25 @@ def _cap_blob_parallelism(cls, workers: int, 
blob_parallelism: int) -> int:
         return max(1, cls._MAX_TOTAL_BLOB_WORKERS // workers)
 
     def _should_run_parallel(
-        self,
-        splits: List[Split],
-        effective: int,
+            self,
+            splits: List[Split],
+            effective: int,
     ) -> bool:
         """Decide whether to take the parallel read path.
 
         ``effective == 1`` falls back to the serial path (no thread pool
         overhead, no behavior change). A single split is never
         parallelized since there is nothing to fan out across.
         """
-        deferred_limit_may_prune = (
-            self.limit is not None
-            and self._deferred_blob_fields
-            and not self._limit_covers_all_splits(splits)
-        )
+        deferred_limit_may_prune = self._deferred_blob_limit_may_prune(splits)
         return (effective >= 2 and len(splits) >= 2
                 and not deferred_limit_may_prune)
 
+    def _deferred_blob_limit_may_prune(self, splits: List[Split]) -> bool:
+        return (self.limit is not None
+                and self._deferred_blob_fields

Review Comment:
   Fixed in b606d8a426. Native LIMIT eligibility now also detects configured 
blob-descriptor-field/blob-view-field payloads and falls back to the Python 
reader when the limit may prune rows. Added a real-binding regression with a 
discarded missing descriptor payload and verified only the retained payload is 
opened.



##########
paimon-python/pypaimon/read/table_read.py:
##########
@@ -239,6 +262,10 @@ def _try_to_pad_batch_by_schema(batch: 
pyarrow.RecordBatch, target_schema):
         for field in target_schema:
             if field.name in batch.schema.names:
                 col = batch.column(field.name)
+                if col.type != field.type:

Review Comment:
   Fixed in b606d8a426. Native reads now fall back for precision-zero 
timestamp[s] fields, including TIMESTAMP_LTZ(0) with UTC, because the current 
Rust binding emits milliseconds for these precisions. The schema check is 
recursive for nested Arrow types, and real-binding regression tests cover both 
timestamp variants.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to