MisterRaindrop opened a new issue, #2008:
URL: https://github.com/apache/cloudberry/issues/2008

   Everything between the DDL skeleton (#1842) and a lake table that can be 
written and read from SQL. Design: #1683. This issue is the execution view.
   
   Landed: #1842, the skeleton (AM, catalog/volume FDWs, stub engine, DDL 
guards); #1951, the Parquet format layer (Arrow, field ids, tracked memory 
pool, type rules).
   
   ### Twelve PRs, three waves
   
   ```
    wave 1                           wave 2                            wave 3
    A  S3 storage ────────────────────────────────────────┐
    B0 proto ─┬► B1 agent framework ─┬► B3 Iceberg ops ─► B4 Builtin ─┼► C 
write ─┬► E merge-on-read + DML
              └► B2 engine client ───┘        ├► B5 Polaris            ├► D 
read ──┘
                                              ├► B6 Hive Metastore
                                              └► B7 Hadoop
   ```
   
   C and D need A, B2 and B4; with the Builtin catalog they run in CI against 
MinIO and the agent alone. B5–B7 need B3 and run alongside C and D.
   
   - [ ] #A   Storage I/O over object storage (S3) — first
   - [ ] #B0  The gRPC contract: proto files
   - [ ] #B1  datalake_agent: service framework (Java)
   - [ ] #B2  datalake_fdw: the agent engine client
   - [ ] #B3  datalake_agent: Iceberg operations by metadata location
   - [ ] #B4  datalake_fdw: Builtin catalog
   - [ ] #B5  datalake_agent: Polaris (REST) catalog
   - [ ] #B6  datalake_agent: Hive Metastore catalog
   - [ ] #B7  datalake_agent: Hadoop (filesystem) catalog
   - [ ] #C   Write path: INSERT to data files and a committed snapshot
   - [ ] #D   Read path: planned fragments, CustomScan, projection and pruning
   - [ ] #E   Merge-on-read and DML
   
   Attached to the PR they belong with: #1988 NUMERIC (C), #1990 timestamp 
units (C), #1989 name mapping (D). Later, blocking nothing: the optional 
`datalake_proxy` bgworker (#1683 §5.4).
   
   ### Rules
   
   - PRs target `main` and are squash-merged; the description is the commit 
message (#1951 layout).
   - Every PR leaves both components shippable: an unfinished path refuses with 
a clear error. The agent answers UNIMPLEMENTED for RPCs it does not carry yet, 
as the stub engine does.
   - The proto files and the shared headers (`format/format.h`, 
`meta/iceberg_meta_engine.h`, `common/*.h`) change only in their own small PR.
   - Nothing generated is committed; protoc runs at build time on both sides.
   - At most two PRs in review at once; the next stays a draft.
   - Each PR brings its tests: a smoke category on the C side, `mvn -B verify` 
on the Java side, a CI job for any service it needs.
   - Milestone `datalake_fdw: Iceberg read/write MVP`, label `datalake`.
   
   Out of scope: the FDW raw-file path (#1683 §2.3), the CI build image (devops 
repo), HDFS storage, execution-engine work.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to