zy-kkk opened a new issue, #68036:
URL: https://github.com/apache/doris/issues/68036
### Search before asking
- [x] I had searched in the issues and found no similar issues.
### Description
Track complete support for Iceberg REST server-side scan planning in Doris.
When a table requires `server` mode, Doris should obtain the plan through the
SDK and execute the tasks returned by the server.
Doris already uses Iceberg Java SDK 1.11.0, which provides `RESTTable` and
`RESTTableScan`. However, Doris snapshot wrappers currently rebuild `RESTTable`
as plain `BaseTable`, losing its planning behavior and bypassing the existing
unsupported-mode guard. Bug #68035 and its fix track restoring rejection. This
tracking issue covers the subsequent implementation of supported server-planned
reads.
#### Delivery sequence
| Stage | Deliverable | Issue completion |
|---|---|---|
| PR 1: restore rejection | Reject real REST server-planned tables before
local manifest access through both statement and cache snapshot paths; retain
normal client-planned behavior | Close #68035; keep this feature open |
| PR 2: implement support | REST planning, task execution, residuals,
scan-scoped credentials, lifecycle handling and regression coverage before
enabling supported reads | Close this feature when the acceptance criteria
below pass |
Two PRs are the initial plan. PR 2 may be split into preparatory
refactoring, test infrastructure and documentation changes if needed.
Intermediate versions must continue rejecting unsupported server-planned reads:
do not enable queries before residual enforcement and credential handling are
ready.
#### Technical approach (proposed)
1. **Retain table capabilities while pinning metadata.** Identify
server-planned tables before snapshot wrapping can discard their behavior.
Design a read view that keeps the selected metadata generation and REST scan
semantics together, with the same catalog/FileIO ownership guarantees. Do not
merely substitute an unfrozen mutable table.
2. **Branch before local manifest access.** Route supported reads through
the SDK REST scan. Bypass manifest-based estimates, COUNT shortcuts and
manifest-cache planning in both eager and streaming paths. Keep snapshot/schema
and projection handling consistent with the requested read.
3. **Translate the complete task semantics.** Reuse data/delete range
construction where appropriate, but carry each task's residual to BE and AND it
with the SQL predicates. The encoding and required columns must be verified,
including filter-only columns and unsupported expressions.
4. **Bind credentials after planning.** Read `scan.fileIO().get()` after
`planFiles()` supplies the plan. Resolve any scan-scoped storage credentials
and make them available before BE scan properties/ranges are finalized.
Preserve the SDK's no-scan-credentials fallback; do not cache a query's
credentials for other queries.
5. **Own the task lifecycle.** Keep iterables alive while their tasks are
consumed, and close them on completion, cancellation or failure. Streaming
implementations must transfer ownership to the split source rather than close
it when the planning method returns.
```java
// Illustrative planning sequence, not a patch or a complete implementation.
TableScan scan = table.newScan(); // Must retain RESTTableScan behavior.
// Apply the selected snapshot/schema, projection and query filter before
planning.
try (CloseableIterable<FileScanTask> tasks = scan.planFiles()) {
FileIO scanIO = scan.fileIO().get(); // Available after planning.
for (FileScanTask task : tasks) {
// Convert data/delete files, enforce task.residual(), and bind
scanIO credentials.
}
}
```
Two existing integration points illustrate where the new branch is required.
`streamingSplitEstimate` currently reads manifests before task enumeration:
https://github.com/apache/doris/blob/28577df9f6df3446f7fd1d0c90032fb7d85c7733/fe/fe-connector/fe-connector-iceberg/src/main/java/org/apache/doris/connector/iceberg/IcebergScanPlanProvider.java#L494-L499
`planFileScanTask` currently chooses between SDK enumeration and local
manifest-cache planning:
https://github.com/apache/doris/blob/28577df9f6df3446f7fd1d0c90032fb7d85c7733/fe/fe-connector/fe-connector-iceberg/src/main/java/org/apache/doris/connector/iceberg/IcebergScanPlanProvider.java#L2698-L2711
This is an implementation direction, not a settled API design. Verify
scan-planning versus `getScanNodeProperties` ordering and the SPI/BE
representation of per-file residuals before deciding whether framework or BE
changes are needed. The existing guard must remain effective until these paths
are complete.
#### Checklist
**Prerequisite**
- [ ] Merge the guard fix with real `RESTTable` coverage for both snapshot
paths.
- [ ] Preserve metadata snapshot isolation, cache leases and resource
ownership.
**Planning and execution**
- [ ] Honor client `scan-planning-mode` and server table overrides, with
server configuration taking precedence and client remaining the default.
- [ ] Use SDK REST planning for synchronous completion, asynchronous polling
and batched task retrieval.
- [ ] Bypass local manifest caching, manifest-based split estimation and
COUNT optimizations in server mode; do not require local manifest access before
submitting the plan.
- [ ] Submit query filters, projected columns and the selected snapshot;
handle historical schemas, time travel and branch/tag reads correctly.
- [ ] Convert planned data tasks and delete files into Doris scan tasks,
covering position deletes, equality deletes and applicable deletion vectors.
- [ ] Enforce each task's residual together with user predicates. Reject
residuals that cannot be represented or executed rather than ignoring them.
- [ ] Propagate plan-provided scan-scoped credentials to data readers after
planning, following SDK fallback semantics and avoiding cross-query credential
reuse.
**Lifecycle and observability**
- [ ] Close task iterators and release SDK/server resources on cancellation,
failure and completion; verify synchronous and asynchronous paths.
- [ ] Report timeouts, server failures and missing endpoints without
automatically falling back to local planning.
- [ ] Expose planning mode, relevant timings, task counts and available plan
identifiers in EXPLAIN/profile without exposing credentials.
- [ ] Document configuration, server overrides, supported operations and
limitations. Decide whether Doris-prefixed aliases are useful during
implementation; raw properties already reach the SDK.
#### Acceptance criteria
| Scenario | Expected supported behavior |
|---|---|
| Normal client mode | Existing results and features remain correct |
| Client requests server mode or server forces it | REST planning endpoints
are used and returned tasks are executed |
| Server adds `id <= 2`; SQL has no WHERE clause | Match the official
reader: return 1 and 2, not 3 and 4 |
| Client cannot read manifests but can read data; REST can read metadata |
Query succeeds without FE manifest reads |
| Data access requires credentials returned by the plan | Use scan-scoped
credentials successfully without credential leakage across queries |
| Unsupported residual, planning failure or timeout | Explicit failure, no
silent local fallback |
| Time travel, schema evolution, branch/tag and delete files | Match
reference results within the supported scope |
| Cancellation, closure, asynchronous and batched plans | Correct query
termination and resource release |
Unit, integration and Doris SQL regression tests should cover these
behaviors together. For larger tables, also compare FE planning time and
manifest access; performance improvements should be measured rather than
assumed.
#### Existing validation baseline
A Docker REST fixture and MinIO environment, a local policy proxy, the
official Iceberg reader and Doris already provide these controls:
- The official reader reads actual Parquet data and applies server
residuals, returning rows 1 and 2. Current Doris returns 1 through 4 without
requesting `/plan`.
- With client metadata access denied, SDK server planning and data-file
access succeed; current Doris fails with 403 while reading the manifest list
locally.
- A real `RESTTable` triggers the existing guard before snapshot wrapping
but not afterwards, demonstrated by method-level controls.
These provide acceptance-test baselines. Asynchronous planning, scan-scoped
credential vending and delete-file cases still need additional validation.
#### Scope and design questions
The first supported scope is REST data-table reads. Remote signing,
BE-distributed task retrieval, automatic fallback, server-planned writes and
server-planned metadata-table queries are excluded. Unsupported paths must
remain explicit rather than silently losing their table semantics.
Before implementation, verify the ordering of scan planning and
`getScanNodeProperties`, and how per-file residuals can reach BE execution.
Whether connector SPI or BE changes are necessary depends on those findings;
this proposal does not assume an FE-only implementation.
### Use case
- Move large-table manifest reads and planning to a REST service that can
use its own caches or indexes.
- Query tables where clients cannot read metadata files directly but can
execute planned data tasks.
- Execute file selections and row filters supplied by the server.
- Read data with temporary credentials issued for an individual scan.
### Related issues
- Prerequisite: #68035, snapshot wrappers bypassing the server-planning
guard.
- Related PR: #67687, SDK upgrade and the original unsupported-mode guard.
### Are you willing to submit PR?
- [x] Yes I am willing to submit a PR!
### Code of Conduct
- [x] I agree to follow this project's Code of Conduct.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]