xiangfu0 opened a new pull request, #19708: URL: https://github.com/apache/pinot/pull/19708
The intermittent `Query check FAILED: lookup() join against dimBaseballTeams` in the Pinot Quickstart workflow is a real bug. Servers that share a JVM also shared one `DimensionTableDataManager`, and a server that initialized it late wiped the lookup table another server had already loaded. `DimensionTableDataManager.createInstanceByTableName` returned a JVM-wide singleton per table. The BATCH quickstart runs three servers in one JVM (it took over the 3-server JOIN quickstart in #19090), so all three servers got the same object and each called `init()` on it. `init()` rebinds the manager to that server's data directory, Helix manager and download semaphore, and `doInit()` resets the lookup table to empty. In the failing jobs ([1](https://github.com/apache/pinot/actions/runs/36521371101/job/109254761712), [2](https://github.com/apache/pinot/actions/runs/36534829172/job/109296373549), [3](https://github.com/apache/pinot/actions/runs/36572045660/job/109418184193)), one server created the manager, downloaded the segment and logged `Successfully loaded lookup table`. The other two then re-initialized the same object, which emptied the lookup table. Their `addOnlineSegment` found the segment already in the shared segment map (`has CRC ... same as before, not replacing it`), so nothing reloaded it. `lookup()` then failed with `Dimension table is not populated` for the whole five-minute window. The multi-stage join on the same table passed because it reads the segments directly. Runs pass when all three `init()` calls land before the first load finishes, as in [this passing run](https://github.com/apache/pinot/actions/runs/36570108108/job/109411700992). The concurrent `init()` calls also show up as `Unable to create temp resources directory` and `Maximum permit count exceeded` errors in these logs, even when the check passes. With this change every server gets its own manager. A manager registers for lookups once `init()` succeeds and deregisters itself on shutdown. When several servers in one JVM host the table, `getInstanceByTableName` prefers one that has loaded its lookup table. A production server runs in its own JVM, so it still has exactly one manager per table. The only difference there is that lookups no longer see a manager whose `init()` failed or that is shutting down; they get `Dimension table does not exist` instead of an NPE or a half-built table. `createInstanceByTableName` keeps its signature for existing `TableDataManagerProvider` implementations, but it now always returns a new instance. Testing: - Three new tests in `DimensionTableDataManagerTest` cover three cases: a server initializing after another has loaded the segment, lookups preferring a populated manager, and a failed `init()` registering nothing. All three fail against the old class, the first with the same "not populated" symptom. - `DimensionTableDataManagerTest`, `LookupTransformFunctionTest`, `LookupJoinOperatorTest`, `ResourceBasedQueriesTest` and `DimensionTableIntegrationTest` pass. - I built the bin-dist on JDK 25 and ran the BATCH quickstart five times with the CI script's checks plus a multi-stage lookup join. Every run passed and `lookup()` succeeded on the first poll. Each server downloaded the segment into its own data directory and loaded its own lookup table, with none of the errors above. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
