shyjsarah opened a new pull request, #881:
URL: https://github.com/apache/paimon-rust/pull/881

   ### Purpose
   
   Closes #872. Supersedes #873 with a fix at the catalog boundary instead of 
changing Tokio blocking behavior.
   
   DataFusion's metadata discovery APIs are synchronous, while Paimon catalog 
and external table-engine resolution are asynchronous. Calling those async 
paths from `schema_names`, `schema`, `table_names`, or `table_exist` can block 
a DataFusion worker that the query itself needs, which is exposed by filtered 
`information_schema` queries with repartitioning.
   
   ### Brief change log
   
   - Build database, table, and view metadata snapshots asynchronously when 
catalog/schema providers are initialized or explicitly refreshed.
   - Make synchronous catalog discovery callbacks read only the in-memory 
snapshot; concrete table loading remains asynchronous.
   - Refresh stale metadata before `SHOW`/`information_schema` queries and 
unknown-table planning, and refresh best-effort after metadata-changing DDL.
   - Preserve the last good snapshot when a refresh fails.
   - Initialize and expose metadata refresh through the Python `PaimonCatalog` 
wrapper.
   - Keep catalog-declared external tables discoverable without synchronously 
invoking their engine resolver.
   
   ### Tests
   
   - `cargo fmt --all --check`
   - `cargo clippy -p paimon-datafusion --lib --tests -- -D warnings`
   - `cargo test -p paimon-datafusion --test sql_context_tests` (60 passed)
   - `cargo test -p paimon-datafusion --test table_type_routing` (38 passed)
   - Targeted Python catalog-provider regressions (2 passed)
   - `cargo test -p paimon-datafusion --lib`: 352 passed; the remaining 7 
require the absent shared `partitioned_log_table` / 
`multi_partitioned_log_table` fixtures
   
   ### API and Format
   
   Adds asynchronous `try_new` and `refresh_metadata` APIs to 
`PaimonCatalogProvider` and `PaimonSchemaProvider`, plus 
`PaimonCatalog.refresh_metadata()` in Python. Existing synchronous constructors 
remain available and start with an empty snapshot, so direct callers should 
initialize or refresh before registration.
   
   No storage format changes.
   
   ### Documentation
   
   Updated provider API documentation to describe snapshot initialization and 
refresh behavior.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to