DanielLeens commented on issue #7937:
URL: https://github.com/apache/seatunnel/issues/7937#issuecomment-5083135821
## Dynamic Lookup STIP candidate and Phase 0 boundary
I am posting the reviewed architecture candidate here so that the umbrella
request has a concrete proposal to discuss. This comment does not claim an
assigned STIP number or formal Apache approval.
**Current status:** Proposed architecture candidate. Only an M0 / Phase 0
infrastructure spike is ready to be explored.
**Not approved or implemented:** the M1 lookup baseline, Experimental or GA
support, production use, or an Accepted/Final architecture decision.
Architecture snapshot SHA-256:
`81465412d02651e88e3ea5f763d8eb285369b82e422c7ade82f87ab6b3403ff4`.
### Narrow target scenario
The proposal is intentionally narrower than a general join engine:
- streaming dimension enrichment only;
- one append-only Fact input, with Kafka as the first capability target;
- one single-table Dimension bootstrap/changelog input, with MySQL CDC as
the first capability target;
- primary-key equality lookup;
- latest processing state for future Fact records, without retroactive
correction of already emitted results;
- same-parallelism restore in the first version;
- primary-key changes, unsupported schema changes, shared source gates, and
unsupported topologies fail fast.
The two explicit inputs are:
- port `0`: Fact;
- port `1`: Dimension bootstrap/changelog.
This is not a general relational join, union, deduplication, windowing,
temporal-table engine, or arbitrary multi-stream computation proposal. It also
does not reuse the existing single-input Transform contract, because the two
inputs have different schemas, lifecycle rules, checkpoint ownership, and
recovery order.
### Required engine architecture
The current logical graph can describe multiple upstream actions, but the
execution path does not preserve a true multi-input operator. A multi-input
target is split into source-rooted pipelines, and the physical flow is rebuilt
once per source. That duplicates the target action and loses input-port and
channel identity.
The proposed chain is:
```text
DynamicLookupAction
-> PortAwareLogicalEdge
-> PortAwareExecutionEdge
-> one Pipeline target vertex
-> one physical multi-input target per target subtask
-> deployment input-port/channel descriptors
-> MultiInputTask
```
The same `operatorUid` also owns exactly one control-plane chain:
```text
DynamicLookupAction
-> DynamicLookupCoordinatorAction
-> DynamicLookupCoordinatorTask
-> CoordinatorStateKey(operatorUid)
-> Physical / Deployment / Checkpoint plans
```
The complete architecture additionally defines versioned channel identities,
attempt fencing, per-channel barriers, managed state handles, checkpoint
terminal CAS, source restore gates, bootstrap ordering, and bounded resource
admission. Those protocols are later phase gates and are not being claimed by
the first Draft PR.
### Phase 0 sequence
Phase 0 must remain serial and independently reviewable:
1. PR-1: port-aware multi-input physical topology;
2. PR-2: HASH exchange and canonical key encoding;
3. PR-3: per-channel barrier/alignment state machine;
4. PR-4: managed state, handles, restore, and checkpoint terminal lifecycle;
5. PR-5: dimension bootstrap, Fact gate, leases, and attempt fencing.
M1 lookup semantics cannot start until those gates, compatibility fixtures,
fault injection, and resource bounds have evidence and receive another
maintainer review.
### First Draft PR scope
The first Draft PR is limited to a non-user-reachable PR-1
planning/deployment skeleton:
1. add a versioned port-aware logical edge under a new serialization class
ID;
2. preserve edge identity, target input port, and the versioned exchange
envelope through Logical, Execution, Pipeline, Physical, Deployment, and
Checkpoint plans;
3. keep one target action/vertex instead of duplicating it for each source
root;
4. create immutable logical and physical channel descriptors;
5. create one multi-input task shell per target subtask and fail fast if
someone tries to run it before the later protocols exist;
6. create one operator-scoped coordinator task and checkpoint key;
7. keep existing single-input, Union, multiple-Source, and multiple-Sink
graphs on their legacy path.
PR-1 does **not** add user-facing Dynamic Lookup syntax and does not claim
working lookup execution. It does not implement HASH routing, business-record
transport, barrier alignment, managed lookup state, source restore gates,
bootstrap commands, or production fencing.
### Compatibility invariants
- The existing `LogicalEdge` class ID and exact two-`long` payload remain
frozen.
- The new edge uses a separate tagged class ID and explicit format version.
- Legacy and port-aware edges are symmetrically unequal.
- New edge identity includes edge ID, both endpoints, target input port, and
canonical exchange bytes.
- Reusing an edge ID with a different descriptor fails fast.
- Pure legacy graphs retain their current pipeline split and task topology.
- A mixed-version deployment capability handshake is still an open PR-1 gate
and will not be represented as complete by DTOs alone.
### Requested review
The initial review should focus on whether this narrow scenario and Phase 0
split are acceptable, especially:
- whether the feature must remain separate from general Transform and SQL
join semantics;
- whether the old `LogicalEdge` compatibility boundary is sufficient;
- whether PR-1 should be split further before real deployment-attempt
fencing is added;
- whether the Fact/Dimension source ownership restrictions are acceptable
for the first version.
The implementation PR will remain Draft and will use `Refs #7937`; it will
not close this umbrella issue.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]