xiangfu0 opened a new pull request, #19708:
URL: https://github.com/apache/pinot/pull/19708

   The intermittent `Query check FAILED: lookup() join against 
dimBaseballTeams` in the Pinot Quickstart workflow is a real bug. Servers that 
share a JVM also shared one `DimensionTableDataManager`, and a server that 
initialized it late wiped the lookup table another server had already loaded.
   
   `DimensionTableDataManager.createInstanceByTableName` returned a JVM-wide 
singleton per table. The BATCH quickstart runs three servers in one JVM (it 
took over the 3-server JOIN quickstart in #19090), so all three servers got the 
same object and each called `init()` on it. `init()` rebinds the manager to 
that server's data directory, Helix manager and download semaphore, and 
`doInit()` resets the lookup table to empty.
   
   In the failing jobs 
([1](https://github.com/apache/pinot/actions/runs/36521371101/job/109254761712),
 
[2](https://github.com/apache/pinot/actions/runs/36534829172/job/109296373549), 
[3](https://github.com/apache/pinot/actions/runs/36572045660/job/109418184193)),
 one server created the manager, downloaded the segment and logged 
`Successfully loaded lookup table`. The other two then re-initialized the same 
object, which emptied the lookup table. Their `addOnlineSegment` found the 
segment already in the shared segment map (`has CRC ... same as before, not 
replacing it`), so nothing reloaded it. `lookup()` then failed with `Dimension 
table is not populated` for the whole five-minute window. The multi-stage join 
on the same table passed because it reads the segments directly.
   
   Runs pass when all three `init()` calls land before the first load finishes, 
as in [this passing 
run](https://github.com/apache/pinot/actions/runs/36570108108/job/109411700992).
 The concurrent `init()` calls also show up as `Unable to create temp resources 
directory` and `Maximum permit count exceeded` errors in these logs, even when 
the check passes.
   
   With this change every server gets its own manager. A manager registers for 
lookups once `init()` succeeds and deregisters itself on shutdown. When several 
servers in one JVM host the table, `getInstanceByTableName` prefers one that 
has loaded its lookup table.
   
   A production server runs in its own JVM, so it still has exactly one manager 
per table. The only difference there is that lookups no longer see a manager 
whose `init()` failed or that is shutting down; they get `Dimension table does 
not exist` instead of an NPE or a half-built table. `createInstanceByTableName` 
keeps its signature for existing `TableDataManagerProvider` implementations, 
but it now always returns a new instance.
   
   Testing:
   - Three new tests in `DimensionTableDataManagerTest` cover three cases: a 
server initializing after another has loaded the segment, lookups preferring a 
populated manager, and a failed `init()` registering nothing. All three fail 
against the old class, the first with the same "not populated" symptom.
   - `DimensionTableDataManagerTest`, `LookupTransformFunctionTest`, 
`LookupJoinOperatorTest`, `ResourceBasedQueriesTest` and 
`DimensionTableIntegrationTest` pass.
   - I built the bin-dist on JDK 25 and ran the BATCH quickstart five times 
with the CI script's checks plus a multi-stage lookup join. Every run passed 
and `lookup()` succeeded on the first poll. Each server downloaded the segment 
into its own data directory and loaded its own lookup table, with none of the 
errors above.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to