[ 
https://issues.apache.org/jira/browse/UNOMI-967?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Serge Huber updated UNOMI-967:
------------------------------
    Component/s:     (was: unomi-3.1.0)

> Clustered scheduler locks and recovery are still unsafe under node races
> ------------------------------------------------------------------------
>
>                 Key: UNOMI-967
>                 URL: https://issues.apache.org/jira/browse/UNOMI-967
>             Project: Apache Unomi
>          Issue Type: Bug
>          Components: unomi(-core)
>    Affects Versions: unomi-3.1.0
>            Reporter: Serge Huber
>            Assignee: Serge Huber
>            Priority: Major
>             Fix For: unomi-3.1.0
>
>
> h2. Summary
> In a multi-node Unomi cluster, scheduled tasks can still lose locks 
> incorrectly, run on the wrong node after a crash, or stay stuck after a node 
> stops. Operators cannot rely on exclusive tasks and recovery without risking 
> duplicate or stranded work.
> h2. What is going wrong today
> * A node can clear or overwrite another node's lock during release or 
> completion races
> * Crash recovery can dispatch work from non-executor nodes, or leave tasks 
> stranded after shutdown
> * Exclusive (non-parallel) tasks default and type-level checks are not strong 
> enough for operators who expect one active run
> * Lock verification false-negatives after CAS can leave locks held or drop a 
> peer's claim
> * Unit coverage around managers, multi-node isolation, and these races is too 
> thin to catch regressions
> h2. Impact
> * Duplicate or missed task execution in production clusters
> * Tasks stuck in RUNNING/CRASHED until manual intervention or lock timeout
> * Hard-to-reproduce CI flakes when scheduler races are involved
> h2. Who is affected
> * Operators running Unomi with more than one scheduler/executor node
> * Contributors changing the scheduler who lack regression tests for cluster 
> races
> h2. What good would look like
> * Exclusive tasks do not steal or wipe a peer's valid lock
> * Terminal and crash transitions only commit when the node still owns the work
> * Non-executor nodes can mark crashes but do not dispatch recovery unless the 
> task is meant to run on all nodes
> * Shutdown after prepare does not leave orphan locks
> * Solid unit tests cover lock, execution, recovery, provider shutdown, and 
> multi-node isolation races
> * Scheduler manual documents states, locks, and ES/OpenSearch OCC 
> expectations clearly



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to