yuqi1129 opened a new issue, #13581: URL: https://github.com/apache/gravitino/issues/13581
## What would you like to be improved? On main `ab0785324c5cb391ae63e4501d349f4336b8d3aa`, the default change-log poller reads at most 2,000 records and waits three seconds after each cycle. Its maximum consumption rate is therefore below 667 records/s, including when a backlog remains. In a local two-JVM/MySQL 8.0.35 experiment, 32 independent-object writers sustained about 1,165 successful updates/s. Eight separately warmed metalakes were each updated once on A and read from B. Five were still stale at the 15-second test deadline; all eight acknowledged values were verified directly in MySQL. Poll failures were zero and the lag gauge grew substantially. With only `gravitino.entityChangeLog.pollIntervalSecs=1` changed after restarting both nodes with cold caches, all eight trials succeeded, with maximum visibility delay about 1.26 seconds. The 15-second deadline is an experiment threshold, not a claimed product SLA. This is propagation-capacity evidence, not evidence of permanent data loss. ## How should we improve? When a full batch is fetched, continue draining with a bounded work/time budget rather than sleeping for the full normal interval. Keep the ordinary delay for an empty or caught-up poll and preserve failure backoff. Consider a configurable batch size and document propagation capacity in records/s. Add backlog and distinct-entity visibility tests under sustained writes. Reproduction: the `main-occ-cache-20260928` benchmark's `probe_backlog.compare_distinct()` uses two 512 MiB Java 17 JVMs, simple authentication, authorization disabled, a dedicated four-CPU MySQL container, 32 independent writers and eight distinct marker entities. Evidence: `probe-visibility-distinct-poll3s.json`, `probe-visibility-distinct-poll1s.json`, the corresponding durable TSVs, SQL snapshots and Prometheus samples. Minimal reproduction: start two main servers against the same MySQL with cache enabled and default polling; create 64 independent metalakes plus eight marker metalakes; run 32 concurrent property-update clients alternating nodes and sustain more than 667 updates/s; for each marker, GET on B to warm it, PUT a unique property value on A once, then poll B every 50 ms for up to 15 seconds while writers continue. Verify the acknowledged marker values directly in `metalake_meta`. Restart both servers and repeat with the one-second interval. Related prior tracking: [#12377](https://github.com/apache/gravitino/issues/12377) already proposed draining more than one batch per poll cycle. This measured follow-up shows the fixed-batch capacity limitation remains on the tested main SHA; it does not claim the idea is new. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
