Cribbee opened a new issue, #10326:
URL: https://github.com/apache/paimon/issues/10326

   ### Search before asking
   
   - [x] I searched existing issues. The earlier request in #3440 is related 
and is referenced below.
   
   ### Motivation
   
   Paimon already provides a shared Query Service for fixed-bucket primary-key 
tables. Its current backend loads lookup files on demand. Local FULL lookup 
materializes table state inside each lookup job, with shuffle lookup available 
to distribute that state within a job.
   
   It would be useful to offer full-cache materialization in Query Service as 
well: multiple lookup jobs could use the same preloaded dimension state, with 
each service executor holding only its assigned partition/bucket shards. Users 
could explicitly choose to complete materialization before the service is 
registered, rather than relying on queries to trigger lookup-file loading.
   
   This is an additional backend choice for the existing service. The expected 
tradeoff is more bootstrap work and local storage in exchange for preloaded 
lookup state. PARTIAL remains useful when queries access only a small portion 
of a table. Existing optimizations, including reuse of persisted remote lookup 
files where supported, should be considered when comparing the two modes.
   
   There is related historical interest in #3440, which requested full loading 
and ongoing synchronization in Query Service. Its reported performance figures 
describe that older environment; they are not evidence about the current 
implementation. I do not have comparative benchmark data to include, so this 
proposal makes no throughput or latency improvement claim.
   
   ### Solution
   
   Add an opt-in `query-service.cache=FULL` backend, keeping `PARTIAL` as the 
default:
   
   - Reuse the existing full-cache lookup implementation and assign data using 
the same partition/bucket routing as Query Service requests.
   - Bootstrap each executor's assigned shards before registering the complete 
service, then refresh them incrementally without requiring client lookups to 
trigger refresh.
   - Preserve the existing RPC protocol and service discovery. Lookup clients 
continue to use `lookup.cache=AUTO`; client-side `lookup.cache=FULL` continues 
to select a local full cache.
   - Retain the existing fixed-bucket primary-key table requirements and 
storage format.
   
   The initial scope is deliberately limited. Executors materialize complete 
rows for their assigned shards, including all partitions. Bootstrap and rebuild 
require local disk and time; refresh can delay requests within an executor. 
Different executors may temporarily serve different snapshot versions. The 
draft implementation rebuilds after overwrites or expired scan snapshots and 
stops serving after a refresh failure instead of continuing with partially 
updated state. These lifecycle and consistency boundaries are part of the 
design discussion.
   
   ### Anything else?
   
   - Draft implementation: #10325. Focused local tests and Flink SQL 
integration tests have passed; real-environment end-to-end validation is still 
pending.
   - Background: [PIP-10: Introduce Paimon 
QueryService](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/272927133/PIP-10+Introduce+Paimon+QueryService)
 and #3440.
   - I previously discussed the idea with Jingsong Li in the community DingTalk 
group and received feedback that the direction is useful. This issue records 
the proposal for public discussion; the design and scope remain open.
   - Feedback on the proposed scope, configuration, and whether a separate PIP 
is appropriate would be helpful.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR; the current implementation is available as 
draft PR #10325.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to