The GitHub Actions job "Integration Tests for AppVersionRevision" on pekko-management.git/decider has failed. Run started by GitHub user pjfanning (triggered by pjfanning).
Head commit for run: dc8a9f77141226b4c41142ac1dff123e780c91e2 / PJ Fanning <[email protected]> Link the self contact point timeouts, cache resolution, and test both Motivation: Follow-up on the review of this branch. The 30 second wait added here and the 10 second timer in `ClusterBootstrap.ensureSelfContactPoint` are coupled - the wait is only correct while it outlasts the timer that completes the promise - but nothing expressed that. Raising the timer would silently make the decider time out first, replacing the deliberate "'Bootstrap.selfContactPoint' was NOT set" error with a bare TimeoutException from elsewhere. The `lazy val` re-runs its initialiser after a throw, so in the case this change exists for - `start()` never ran, so nothing ever completes the promise - every `canJoinSelf` call blocks a dispatcher thread for the full timeout again. The previous `Duration.Inf` parked one thread once; this parks one per probe, for as long as the contact point stays unset. There was no test. `SelfAwareJoinDeciderSpec` only covers the path where the contact point has already been set, and a 30 second hardcoded timeout cannot be exercised in a test anyway. Modification: Move the timer's duration to `ClusterBootstrap.SelfContactPointTimeout` and derive the decider's wait from it, behind a `protected def` a test can override. Replace the `lazy val` with an `AtomicReference` that caches a resolved value. A failure is deliberately not cached, so a contact point set later is still picked up, but the blocking wait is paid at most once: reaching the timeout means the promise has nothing to complete it, so later callers check the promise without blocking and fail fast until it does complete. Resolve against the promise directly rather than mapping it first, so that the already-completed case needs no dispatcher hop and `value` is meaningful. Result: The two timeouts cannot drift apart. An unset contact point costs one blocking wait rather than one per probe, and is still picked up if it arrives late. Tests: - sbt "management-cluster-bootstrap/test" - 54 succeeded, 0 failed (48 before) - New SelfContactPointResolutionSpec covers resolution, caching, the timeout, the block-once behaviour, late setting after a timeout, and the ordering between the two timeouts - Directional: with the block-once guard removed so that every call waits, "block for the timeout only once, then fail fast" FAILS with "505796658 nanoseconds was not less than 250 milliseconds" - sbt "management-cluster-bootstrap/mimaReportBinaryIssues" - success - sbt "management-cluster-bootstrap/scalafmtCheck" "management-cluster-bootstrap/Test/scalafmtCheck", "+headerCheckAll" - clean References: Refs #908 Report URL: https://github.com/apache/pekko-management/actions/runs/32882516974 With regards, GitHub Actions via GitBox --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
