plusplusjiajia opened a new pull request, #10026:
URL: https://github.com/apache/paimon/pull/10026

   ## Problem
   
   A blob table never reaches Daft's native parquet reader through 
`use_native_reader`: `_reader_routing` rules it out on `_has_blob_columns` 
before it ever looks at `has_auth`, so that refusal is dead code for such a 
table. It reaches the native reader through `_blob_table_native_files`, which 
hands Daft the covering parquet files when no BLOB column is projected, and 
that path never looked at query authorization.
   
   A blob table carrying a column mask or a row filter, read with a scalar-only 
projection, therefore returned raw parquet: both live in pypaimon's reader, 
which was skipped entirely.
   
   ## Change
   
   `_blob_table_native_files` takes `has_auth` and returns `None` for it, the 
way it already does for deletion vectors. The parameter has no default, so a 
new call site has to answer for it rather than silently bypassing the check. 
Both existing call sites, task generation and explain, pass it.
   
   Keeping an authorized split off the native reader exposed a second defect. A 
count or a constant projection pushes `columns=[]`, and the fallback task 
cannot build a Daft record batch from zero arrays. That already broke such a 
query on an ordinary authorized table; the blob table only avoided it by 
bypassing authorization. The fallback now carries the row count in a 
placeholder column and projects it away.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to