joaopamaral opened a new pull request, #44825: URL: https://github.com/apache/superset/pull/44825
### SUMMARY Enable catalog support on `RedshiftEngineSpec`, the same way `PostgresEngineSpec` has it. In Redshift a catalog is a database, and cross-database queries (`database.schema.table`) are supported natively, including for databases created from [datashares](https://docs.aws.amazon.com/redshift/latest/dg/datashare-overview.html). - `supports_catalog`, `supports_dynamic_catalog` and `supports_cross_catalog_queries` are set, so SQL Lab and the dataset editor show the catalog selector (when the connection has `allow_multi_catalog`), and generated SQL uses `catalog.schema.table`. - `adjust_engine_params` sets the selected catalog as the URL database; IAM connections (`redshift_connector`) carry the database in `connect_args["database"]`, which is updated too. - `get_default_catalog` returns the URL database. - `get_catalog_names` lists `SVV_REDSHIFT_DATABASES`. Unlike `pg_database`, it also includes databases created from datashares, which is the main motivation: without catalogs there is no way to browse a datashare consumer database from a Redshift connection in Superset. - Migration `e6d0a676a087` runs `upgrade_catalog_perms(engines={"redshift"})`, as done for PostgreSQL, Trino/Presto/BigQuery/Snowflake and Databricks when catalogs were introduced for them: it adds `catalog_access` permissions, renames `schema_access` permissions to include the default catalog, and populates the `catalog` column on existing Redshift datasets, charts and saved queries. Reflecting the datashare objects themselves (schemas/tables/columns of a consumer database, which are not in `pg_catalog`) is handled on the dialect side in sqlalchemy-redshift/sqlalchemy-redshift#334; this PR works without it for regular databases. ### BEFORE/AFTER SCREENSHOTS OR ANIMATED GIF N/A: the catalog selector is the existing UI, now available for Redshift connections with `allow_multi_catalog` enabled. ### TESTING INSTRUCTIONS 1. Add a Redshift connection (any driver), enable "Allow changing catalogs" (`allow_multi_catalog`) in its advanced settings. 2. In SQL Lab, the catalog dropdown lists the cluster's databases (`SVV_REDSHIFT_DATABASES`); selecting one switches the schema and table lists to that database, and table previews / datasets query it as `catalog.schema.table`. 3. Existing Redshift datasets keep working: after the migration their `catalog` is the connection URL database, and their permissions carry the catalog (`[db].[catalog].[schema]`). 4. Unit tests: `pytest tests/unit_tests/db_engine_specs/test_redshift.py`. ### ADDITIONAL INFORMATION - [ ] Has associated issue: - [ ] Required feature flags: - [ ] Changes UI - [x] Includes DB Migration (follow approval process in [SIP-59](https://github.com/apache/superset/issues/13351)) - [x] Migration is atomic, supports rollback & is backwards-compatible - [x] Confirm DB migration upgrade and downgrade tested - [x] Runtime estimates and downtime expectations provided - [ ] Introduces new feature or API - [ ] Removes existing feature or API Migration notes: same shape as `4081be5b6b74` (Databricks) and `87ffc36f9842` (Trino/Presto/BigQuery/Snowflake). It only touches rows of Redshift connections; runtime is proportional to the number of Redshift datasets/charts, and it connects to each Redshift database to discover non-default catalogs (skipped when the connection fails). No downtime expected. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
