dmitry-chirkov-dremio commented on code in PR #51665:
URL: https://github.com/apache/arrow/pull/51665#discussion_r4213237499


##########
cpp/src/gandiva/tests/concurrent_make_test.cc:
##########
@@ -0,0 +1,212 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements.  See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership.  The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License.  You may obtain a copy of the License at
+//
+//   http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied.  See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+// Regression tests for GH-601: concurrent Projector::Make()/Filter::Make() 
calls for the
+// *same* expression cache key.
+//
+// Before the fix, Make() read the shared object cache once to decide 
`is_cached`, then
+// SetLLVMObjectCache() performed a second, independent read of the same key. 
A thread
+// that saw a miss on the first read and a hit on the second one would both 
pre-load the
+// cached object (defining e.g. "expr_0_0" in its JITDylib) *and* add its own 
freshly
+// compiled IR module defining the same symbol, so LLVM ORC's duplicate-symbol 
detection
+// fired and Make() returned "Failed to add IR module to LLJIT: Duplicate 
definition of
+// symbol".
+//
+
+#include <gtest/gtest.h>
+
+#include <condition_variable>
+#include <memory>
+#include <mutex>
+#include <string>
+#include <thread>
+#include <vector>
+
+#include "arrow/memory_pool.h"
+#include "gandiva/filter.h"
+#include "gandiva/projector.h"
+#include "gandiva/tests/test_util.h"
+#include "gandiva/tree_expr_builder.h"
+
+namespace gandiva {
+
+using arrow::int32;
+
+namespace {
+
+// Releases every waiting thread at once, so the racing Make() calls overlap.
+class StartGate {
+ public:
+  void Wait() {
+    std::unique_lock<std::mutex> lock(mutex_);
+    cv_.wait(lock, [this] { return open_; });
+  }
+
+  void Open() {
+    {
+      std::lock_guard<std::mutex> lock(mutex_);
+      open_ = true;
+    }
+    cv_.notify_all();
+  }
+
+ private:
+  std::mutex mutex_;
+  std::condition_variable cv_;
+  bool open_ = false;
+};
+
+int HardwareConcurrency() {
+  return std::max(1, static_cast<int>(std::thread::hardware_concurrency()));
+}
+
+int NumThreads() { return std::min(32, std::max(8, 2 * 
HardwareConcurrency())); }
+
+// This test only works when the threads are oversubscribed relative to the 
CPUs: the race
+// window is between two reads of the shared cache inside one Make().
+// So if the cap in NumThreads() has pulled the count under 2x, this test
+// cannot fail and must not report success.
+bool HasEnoughOversubscription() { return NumThreads() >= 2 * 
HardwareConcurrency(); }
+
+constexpr int kIterations = 25;
+
+}  // namespace
+
+class TestConcurrentMake : public ::testing::Test {

Review Comment:
   I’m not a fan of this non-deterministic tests. 
   Yes, it does make me feel better than test fails without the corresponding 
change.
   
   How common are the test hooks in the repo?  Can’t really spot anything 
beyond one in async_generator_test.cc
   
   I'd add a static inline std::function test hook, the pattern MergedGenerator 
already uses for error_signaled_hook_for_testing. It would fire in 
Projector::Make and Filter::Make right after the cache lookup, so a 
single-threaded test could insert the entry at that exact point. That 
reproduces the miss-then-hit interleaving deterministically, fails on the old 
code and passes on the fix. The hook would be reset in TearDown(). I'd drop the 
oversubscription skip and keep at most a small best-effort threaded test.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to