The GitHub Actions job "Integration Tests for AppVersionRevision" on 
pekko-management.git/decider has failed.
Run started by GitHub user pjfanning (triggered by pjfanning).

Head commit for run:
dc8a9f77141226b4c41142ac1dff123e780c91e2 / PJ Fanning 
<[email protected]>
Link the self contact point timeouts, cache resolution, and test both

Motivation:
Follow-up on the review of this branch.

The 30 second wait added here and the 10 second timer in
`ClusterBootstrap.ensureSelfContactPoint` are coupled - the wait is only correct
while it outlasts the timer that completes the promise - but nothing expressed
that. Raising the timer would silently make the decider time out first, 
replacing
the deliberate "'Bootstrap.selfContactPoint' was NOT set" error with a bare
TimeoutException from elsewhere.

The `lazy val` re-runs its initialiser after a throw, so in the case this change
exists for - `start()` never ran, so nothing ever completes the promise - every
`canJoinSelf` call blocks a dispatcher thread for the full timeout again. The
previous `Duration.Inf` parked one thread once; this parks one per probe, for as
long as the contact point stays unset.

There was no test. `SelfAwareJoinDeciderSpec` only covers the path where the
contact point has already been set, and a 30 second hardcoded timeout cannot be
exercised in a test anyway.

Modification:
Move the timer's duration to `ClusterBootstrap.SelfContactPointTimeout` and
derive the decider's wait from it, behind a `protected def` a test can override.

Replace the `lazy val` with an `AtomicReference` that caches a resolved value.
A failure is deliberately not cached, so a contact point set later is still
picked up, but the blocking wait is paid at most once: reaching the timeout 
means
the promise has nothing to complete it, so later callers check the promise
without blocking and fail fast until it does complete.

Resolve against the promise directly rather than mapping it first, so that the
already-completed case needs no dispatcher hop and `value` is meaningful.

Result:
The two timeouts cannot drift apart. An unset contact point costs one blocking
wait rather than one per probe, and is still picked up if it arrives late.

Tests:
- sbt "management-cluster-bootstrap/test" - 54 succeeded, 0 failed (48 before)
- New SelfContactPointResolutionSpec covers resolution, caching, the timeout, 
the
  block-once behaviour, late setting after a timeout, and the ordering between
  the two timeouts
- Directional: with the block-once guard removed so that every call waits, 
"block
  for the timeout only once, then fail fast" FAILS with "505796658 nanoseconds
  was not less than 250 milliseconds"
- sbt "management-cluster-bootstrap/mimaReportBinaryIssues" - success
- sbt "management-cluster-bootstrap/scalafmtCheck"
  "management-cluster-bootstrap/Test/scalafmtCheck", "+headerCheckAll" - clean

References:
Refs #908

Report URL: https://github.com/apache/pekko-management/actions/runs/32882516974

With regards,
GitHub Actions via GitBox


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to