GitHub user leborchuk added a comment to the discussion: [Proposal] Apache Cloudberry as PostgreSQL 19 Extensions
Hi! Thank you for the attention to our project ) The article https://github.com/igor-suhorukov/cloudberry/blob/extension_postgresql_19/pg19/doc/greenplum-without-the-fork.md looks more broader and has more interesting details, if you don't mind I'll answer mainly to ideas from you original post rather than technical description here. Because some aspects need discussion at first and only then implementation. 1. Kernel rebasing. It's the major reason why you need community. You cannot offload this task to LLM or whatever else. Of course, LLM will spend you tokens easily and you even could get some looks promising result. But without owning the code you won't get good basement for your further improvements. If you could generate whatever you want why one could need cloudberry at all? Or you just could ask LLM to rebase cloudberry to PG19 (it also complete task in a few days). Ease rebasement - what we constantly bear in mind. And the story we usually share with postgres community. You could get postgres, add here PAX - and now you have analytical data storage format. And the same with other new components. There are not so much of them comparing with original greenplum innovations. But now it's just the matter of time. There will be more and more components, I believe ) It will be great to evolve old approaches and make them extension-based. For example, as you described - make ao and aoco tables as extensions. We could start with some specific functionality and then step-by-step goes to the situation where you have postgresql fork extended with various MPP functions. Not based on postgres our own MPP database. The difference of course in the number of changes in a core part (and conflicts need close attention) and code reusing. The idealistic (I think it's too simple to be true) picture is you have a set of extensions and combining them could easily make your own specific database. How to achieve this - if you are really interested in it and ready to participate, let's continue discussion (good example is https://github.com/apache/cloudberry/discussions/1683 ), volunteer project developers and gradually improve our codebase. LLM will help us, but not replace, we still need to clearly understand what we are doing. 2. Clickhouse benchmarks Kernel rebasing is great but as you have written, we rebase kernel not to the sole kernel version but to achieve something - new functionality or better performance. Good example is clickbench - new kernels could get you better performance. But I need to say modern PG kernel is not enough. PG is just (not fully describe current situation but it's too hard to express it succinctly) not good enough to beat clickhouse in clickbench. Not because clickbench purposely was written by clickhouse developers to beat all other competitors (but because of that, too), but mainly because clickhouse constantly compact (sort) data in background and have many others good improvements. We also could improve our group by facilities and be comparable in some aspects with clickhouse. It's possible, but not easy. if you interested - feel free to create discussion. I believe we could do it. Also we could create our own benchmark to beat all the competitors in it. 3. Not only Postgres You mentioned other Postgres-related projects. We could not only compete with them but also get good approaches from them. For example ORCA (and some other GP components) also is used by polardb - database with shared-everything architecture. We also could get some good approaches from polardb and use it in cloudberry. I'am not good at polardb to say which one - it was just a suggestion. But if you have good candidates - let us know. GitHub link: https://github.com/apache/cloudberry/discussions/2065#discussioncomment-18697011 ---- This is an automatically sent email for [email protected]. To unsubscribe, please send an email to: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
