DanielLeens commented on issue #7937:
URL: https://github.com/apache/seatunnel/issues/7937#issuecomment-5083135821

   ## Dynamic Lookup STIP candidate and Phase 0 boundary
   
   I am posting the reviewed architecture candidate here so that the umbrella 
request has a concrete proposal to discuss. This comment does not claim an 
assigned STIP number or formal Apache approval.
   
   **Current status:** Proposed architecture candidate. Only an M0 / Phase 0 
infrastructure spike is ready to be explored.
   
   **Not approved or implemented:** the M1 lookup baseline, Experimental or GA 
support, production use, or an Accepted/Final architecture decision.
   
   Architecture snapshot SHA-256: 
`81465412d02651e88e3ea5f763d8eb285369b82e422c7ade82f87ab6b3403ff4`.
   
   ### Narrow target scenario
   
   The proposal is intentionally narrower than a general join engine:
   
   - streaming dimension enrichment only;
   - one append-only Fact input, with Kafka as the first capability target;
   - one single-table Dimension bootstrap/changelog input, with MySQL CDC as 
the first capability target;
   - primary-key equality lookup;
   - latest processing state for future Fact records, without retroactive 
correction of already emitted results;
   - same-parallelism restore in the first version;
   - primary-key changes, unsupported schema changes, shared source gates, and 
unsupported topologies fail fast.
   
   The two explicit inputs are:
   
   - port `0`: Fact;
   - port `1`: Dimension bootstrap/changelog.
   
   This is not a general relational join, union, deduplication, windowing, 
temporal-table engine, or arbitrary multi-stream computation proposal. It also 
does not reuse the existing single-input Transform contract, because the two 
inputs have different schemas, lifecycle rules, checkpoint ownership, and 
recovery order.
   
   ### Required engine architecture
   
   The current logical graph can describe multiple upstream actions, but the 
execution path does not preserve a true multi-input operator. A multi-input 
target is split into source-rooted pipelines, and the physical flow is rebuilt 
once per source. That duplicates the target action and loses input-port and 
channel identity.
   
   The proposed chain is:
   
   ```text
   DynamicLookupAction
     -> PortAwareLogicalEdge
     -> PortAwareExecutionEdge
     -> one Pipeline target vertex
     -> one physical multi-input target per target subtask
     -> deployment input-port/channel descriptors
     -> MultiInputTask
   ```
   
   The same `operatorUid` also owns exactly one control-plane chain:
   
   ```text
   DynamicLookupAction
     -> DynamicLookupCoordinatorAction
     -> DynamicLookupCoordinatorTask
     -> CoordinatorStateKey(operatorUid)
     -> Physical / Deployment / Checkpoint plans
   ```
   
   The complete architecture additionally defines versioned channel identities, 
attempt fencing, per-channel barriers, managed state handles, checkpoint 
terminal CAS, source restore gates, bootstrap ordering, and bounded resource 
admission. Those protocols are later phase gates and are not being claimed by 
the first Draft PR.
   
   ### Phase 0 sequence
   
   Phase 0 must remain serial and independently reviewable:
   
   1. PR-1: port-aware multi-input physical topology;
   2. PR-2: HASH exchange and canonical key encoding;
   3. PR-3: per-channel barrier/alignment state machine;
   4. PR-4: managed state, handles, restore, and checkpoint terminal lifecycle;
   5. PR-5: dimension bootstrap, Fact gate, leases, and attempt fencing.
   
   M1 lookup semantics cannot start until those gates, compatibility fixtures, 
fault injection, and resource bounds have evidence and receive another 
maintainer review.
   
   ### First Draft PR scope
   
   The first Draft PR is limited to a non-user-reachable PR-1 
planning/deployment skeleton:
   
   1. add a versioned port-aware logical edge under a new serialization class 
ID;
   2. preserve edge identity, target input port, and the versioned exchange 
envelope through Logical, Execution, Pipeline, Physical, Deployment, and 
Checkpoint plans;
   3. keep one target action/vertex instead of duplicating it for each source 
root;
   4. create immutable logical and physical channel descriptors;
   5. create one multi-input task shell per target subtask and fail fast if 
someone tries to run it before the later protocols exist;
   6. create one operator-scoped coordinator task and checkpoint key;
   7. keep existing single-input, Union, multiple-Source, and multiple-Sink 
graphs on their legacy path.
   
   PR-1 does **not** add user-facing Dynamic Lookup syntax and does not claim 
working lookup execution. It does not implement HASH routing, business-record 
transport, barrier alignment, managed lookup state, source restore gates, 
bootstrap commands, or production fencing.
   
   ### Compatibility invariants
   
   - The existing `LogicalEdge` class ID and exact two-`long` payload remain 
frozen.
   - The new edge uses a separate tagged class ID and explicit format version.
   - Legacy and port-aware edges are symmetrically unequal.
   - New edge identity includes edge ID, both endpoints, target input port, and 
canonical exchange bytes.
   - Reusing an edge ID with a different descriptor fails fast.
   - Pure legacy graphs retain their current pipeline split and task topology.
   - A mixed-version deployment capability handshake is still an open PR-1 gate 
and will not be represented as complete by DTOs alone.
   
   ### Requested review
   
   The initial review should focus on whether this narrow scenario and Phase 0 
split are acceptable, especially:
   
   - whether the feature must remain separate from general Transform and SQL 
join semantics;
   - whether the old `LogicalEdge` compatibility boundary is sufficient;
   - whether PR-1 should be split further before real deployment-attempt 
fencing is added;
   - whether the Fact/Dimension source ownership restrictions are acceptable 
for the first version.
   
   The implementation PR will remain Draft and will use `Refs #7937`; it will 
not close this umbrella issue.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to