reuvenlax opened a new pull request, #40320: URL: https://github.com/apache/beam/pull/40320
An attempt to better parallelize ValidatesRunner turned up a number of problems in our file staging. 1, TestDataflowRunner previously set tempLocation to <tempRoot>/<jobName>/output/results, causing DataflowPipelineOptions.StagingLocationFactory` to derive a unique path for every single test job. Consequently, every test re-hashed ~300+ MB of classpath jars and re-queried/uploaded them to a brand-new GCS staging directory instead of reusing staged artifacts across tests. **Fix**: Defaults stagingLocation to a shared /staging/ directory across jobs in the same test run (while keeping tempLocation isolated per job at <tempRoot>/<jobName>/output/results). Caches SHA-256 file hashes and directory zip artifacts in memory keyed by (path, size, lastModifiedMillis) so unmodified classpath jars are hashed once per worker JVM instead of once per test. Caches resolved DataflowPackage lists per `stagingLocation, filesToStage) when running under TestDataflowRunner so each worker JVM stages/verifies the classpath once. 2. Every call to p.run() would calls PipelineResource.detectClassPathResourcesToStage which used a lot of RAM. Increasing parallelism of tests caused multiple OOMs. ClassGraph.scan(numThreads) opens every JAR and directory on the classpath. However, detect(...) only ever needed the top-level File paths of the classpath elements themselves (.getClasspathFiles()), which ClassGraph resolves directly from the ClassLoader / java.class.path without opening or scanning inside the JARs., parses their ZIP central directories, enumerates every internal resource/class entry inside every JAR, and constructs an in-memory ScanResult graph.Also the ScanResult objects (which contain mmap buffers) were lease. **Fix** calls ClassGraph.getClasspathFiles() directly instead of .scan(1).getClasspathFiles(). Also caches the resolves classpath, and uses a lock to ensure that only on thread resolve the classpath at a time (the other threads should then hit the cache) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
