GitHub user igor-suhorukov created a discussion: [Proposal] Apache Cloudberry as PostgreSQL 19 Extensions
### Proposers _No response_ ### Proposal Status Under Discussion ### Abstract I would like to discuss running Apache Cloudberry on PostgreSQL 19 as a set of extensions on top of 22 dormant core hooks, instead of a forked kernel. A working prototype exists; it is my personal experiment, not an ASF release. Without the modules, the patched server is measurably vanilla PostgreSQL 19: the same PostgreSQL tests pass, the ABI is unchanged, catalogs are byte-identical, and the hooks add less than 0.2% to executed instructions. With them, it is a Greenplum-style MPP engine with plan slices, Motions, two-phase commit and distributed snapshots, running unmodified ORCA and stock PostGIS and pgvector. ORCA plans all 121 TPC-H and TPC-DS queries at SF1 with answers matching DuckDB's, and 691 of Cloudberry's 1,180 test files pass so far. Claude Code wrote the code over 12 days and 629 commits; I set the goals and signed off every change to the PostgreSQL core. ### Motivation Greenplum and Cloudberry have always followed PostgreSQL a few releases behind: Greenplum 6 shipped on PostgreSQL 9.4, Greenplum 7 on 12, Cloudberry 2.x on 14, and `main` reached 16.9 in May 2026. The reason is structural. Cloudberry edits 882 PostgreSQL backend and header files, so every new PostgreSQL major is a large merge. #1095 estimated the 14 → 16 upgrade at 4–5 months, three of them for fixing regression tests, and noted that PostgreSQL 14, the base of Cloudberry 2.x, reaches end of life in 2026. Meanwhile PostgreSQL 17 and 18 have shipped, and 19 is close to release. In #1877, the project named close compatibility with modern PostgreSQL as one of its goals. The prototype tests whether that goal can become cheap to sustain. If Cloudberry runs as extensions, the next PostgreSQL major means rebasing a small patch series instead of merging a fork. Users get a current PostgreSQL, including PostgreSQL 19 features such as eager aggregation and parallel autovacuum. They also get stock extensions: upstream PostGIS 3.7 instead of a PostGIS fork, and pgvector 0.8.6 built unpatched, which is relevant to #2043. The prototype compiles Cloudberry's own sources in place, so it builds on the community's work rather than replacing it. ### Implementation Details: [Greenplum Without the Fork](https://github.com/igor-suhorukov/cloudberry/blob/extension_postgresql_19/pg19/doc/greenplum-without-the-fork.md). Code: [core series](https://github.com/igor-suhorukov/postgres/tree/REL_19_STABLE_CLOUDBERRY), [port](https://github.com/igor-suhorukov/cloudberry/tree/extension_postgresql_19/pg19), [build instructions](https://github.com/igor-suhorukov/cloudberry/tree/extension_postgresql_19/pg19#building). The ground rule: a server with the core patches but without Cloudberry loaded must be indistinguishable from vanilla PostgreSQL 19, and this is measured. - **Core series:** 22 commits, 47 files, +1,578/−41 lines. Existing extension points come first: hooks, table AMs, CustomScan, custom WAL resource managers, background workers, security labels and FDWs. Where they are not enough, a new hook is a NULL-by-default function pointer, a new exported function, a registration list or an off-by-default flag. No existing signature changes, no struct gains a field, and no catalog, WAL, page or protocol format changes. - **Vanilla checks** compare Docker images built from the same PostgreSQL commit with and without the series: - the same 371 `meson test` results; - `abidiff`: 32 symbols added, 0 changed; - byte-identical catalogs, settings, keywords and stored parse trees; - identical query IDs; - interchangeable data directories and streaming replication; - cachegrind: −0.005% to +0.155% executed instructions; - a script that checks every statement added inside an existing function sits behind a hook, registry or flag, with one reviewed exception. - **Cloudberry's tree is untouched.** The whole port is one new directory, `pg19/`. Cloudberry's files compile where they lie, so merges from Cloudberry keep applying, and a file that must change becomes a copy that says what changed. ORCA's 920 core files compile unmodified behind `planner_hook`, and PAX compiles in place with 65 of its 91 files unchanged. - **Modules:** - `gp_core`: dispatcher, Motions, interconnect, 2PC, distributed snapshots, GDD, FTS; - `gp_orca`; - `gp_sql`: Cloudberry syntax rewritten into PostgreSQL 19's grammar; - storage: `gp_ao`, `pax`; - external data: `gp_exttable`, `pxf_fdw`, `gpcloud`, `datalake_fdw`; - `gp_resource`, `gp_security`, `gp_matview`, `gp_task`, `diskquota`; - `gpfts` and gpMgmt's tools from `gpinitsystem` to `gpexpand`. - **What changes for users:** - migration is logical (dump and restore); - settings are renamed `gp.*`, although `gpconfig` accepts the old names; - DDL command tags become standard; - the 20 shared catalogs are replaced by security labels, extension tables and a cluster file; - TDE gives way to volume encryption. ### Rollout/Adoption Plan Nothing here affects Cloudberry 2.x or the 3.0 plans: the prototype lives in my fork and changes no lines in Cloudberry's files. I would like to ask the community for four things: 1. **Feedback on the direction.** Is running Cloudberry as PostgreSQL extensions worth pursuing as a long-term path after 3.0? 2. **A home for `pg19/`.** It is one directory that tracks Cloudberry's sources without editing them. If the direction looks right, would the project consider hosting it, for example as an experimental branch or a separate repository? 3. **Help with the remaining tests.** The goal is to pass Cloudberry's original test suites with minimal changes to the tests. Every skipped test is listed with the statement that stops it, so tests are easy to pick up; `isolation2` and `greenplum_schedule` are next. 4. **A decision on Route B.** Should the port bring over Cloudberry's MPP variant of the PostgreSQL planner (30–45k lines), or keep shrinking ORCA's declines? On the regression suite, the port's own fallback reasons already fell from 563 to 52. A proposed `gp.mpp_planner` switch would allow side-by-side comparison, and I would value the view of the planner experts here. For users, migration would be logical, similar to the cluster-to-cluster copy planned for 2.0 → 3.0 in #1095. Separately, four of the custom hooks (table AM registry, smgr file events, parser hook and OID hook) are general enough to propose to PostgreSQL upstream. Support from Cloudberry developers as their real-world user would strengthen that case. ### Are you willing to submit a PR? - [X] Yes I am willing to submit a PR! GitHub link: https://github.com/apache/cloudberry/discussions/2065 ---- This is an automatically sent email for [email protected]. To unsubscribe, please send an email to: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
