wombatu-kun commented on code in PR #20120:
URL: https://github.com/apache/hudi/pull/20120#discussion_r4120580632
##########
hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/hudi/HoodieSparkSqlWriter.scala:
##########
@@ -971,6 +972,26 @@ class HoodieSparkSqlWriterInternal {
}
}
+ // Spark's own file-based writers invalidate the session cache from
+ // InsertIntoHadoopFsRelationCommand. Hudi writes do not go through that
command, so without
+ // this a cached Hudi table keeps serving the pre-write snapshot with no
signal to the reader.
+ // Matching by path rather than by plan also reaches entries built from a
DataFrame that was
+ // never registered in the catalog, which the refreshTable below cannot
see.
+ //
+ // A failure here must not fail the write. The commit has already
succeeded, and
+ // CacheManager.recacheByCondition drops the matching entries before it
attempts to rebuild
+ // them, so the invalidation has taken effect even when the rebuild
throws. The rebuild does
+ // throw when the cached plan is no longer valid against the table it was
built from: an
+ // overwrite that replaces a partitioned table with a non-partitioned one
leaves the cached
+ // plan holding the old partition schema, and re-optimizing it fails on
the new layout.
+ try {
+ spark.catalog.refreshByPath(basePath.toString)
Review Comment:
When a meta sync fails, `metaSync` throws the collected failure before
reaching this call, so a write that has already committed still leaves the
cached entries stale. Could this block move above the sync block, so the
invalidation no longer depends on the sync succeeding?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]