JingsongLi commented on code in PR #9598:
URL: https://github.com/apache/paimon/pull/9598#discussion_r3969534923


##########
paimon-spark/paimon-spark-common/src/main/scala/org/apache/paimon/spark/PaimonPartitionReader.scala:
##########
@@ -136,3 +149,24 @@ case class PaimonPartitionReader(
     }
   }
 }
+
+private[spark] object BlobDescriptorUtils {
+
+  def createUriReaderFactory(
+      table: Table,
+      readType: RowType,
+      blobAsDescriptor: Boolean): UriReaderFactory = {
+    if (blobAsDescriptor || !hasBlobFileFields(readType)) {
+      null
+    } else {
+      table match {
+        case fileStoreTable: FileStoreTable => 
BlobDescriptorReaderFactory.create(fileStoreTable)

Review Comment:
   [P1] Scope reader rebinding to descriptor fields and preserve managed BLOB 
FileIO
   
   `hasBlobFileFields` also matches ordinary managed `blob-field` columns, so 
this creates an external descriptor factory for them, and 
`BlobDescriptorResolvingRow` subsequently replaces every `BlobRef` reader. For 
a managed primary-key BLOB, `ColumnarRow.getBlob()` already supplies the target 
table's FileIO. Replacing it with a reader rebuilt from catalog options loses 
table-scoped credentials (e.g. `RESTTokenFileIO`), causing default payload 
reads to fail even when `blob-descriptor.source-table` is unset. I reproduced 
this with an isolated table-scoped FileIO: the original BlobRef reads 
successfully, while the wrapped row fails.
   
   The same scope problem introduces an unnecessary source-table dependency 
after materialization. In the existing `Blob: materialize descriptor with 
source table FileIO` test, dropping `blob_source` after the insert makes 
`SELECT picture FROM blob_target` fail with `Failed to load BLOB descriptor 
source table`, although the payload has already been copied into the target. 
Removing the target's `blob-descriptor.source-table` option restores the read.
   
   Please limit both factory initialization and reader replacement to the 
projected `blob-descriptor-field` columns, preserving the original readers for 
managed BLOBs. Gating factory creation alone is insufficient when a projection 
contains both descriptor and managed columns. Please cover table-scoped FileIO 
and reading a materialized copy after source removal in regression tests.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to