Stephen O'Donnell created HDDS-6661:
---------------------------------------

             Summary: TestBackgroundPipelineScrubber.testRun() fails 
intermittently
                 Key: HDDS-6661
                 URL: https://issues.apache.org/jira/browse/HDDS-6661
             Project: Apache Ozone
          Issue Type: Improvement
            Reporter: Stephen O'Donnell


TestBackgroundPipelineScrubber.testRun() fails intermittently for me locally, 
with this trace:

{code}
2022-04-27 17:20:44,888 [main] INFO  pipeline.BackgroundPipelineScrubber 
(BackgroundPipelineScrubber.java:start(123)) - Starting Pipeline Scrubber 
Service.
2022-04-27 17:20:44,890 [main] INFO  pipeline.BackgroundPipelineScrubber 
(BackgroundPipelineScrubber.java:notifyStatusChanged(85)) - Service 
BackgroundPipelineScrubber transitions to RUNNING.
2022-04-27 17:20:47,905 [main] INFO  pipeline.BackgroundPipelineScrubber 
(BackgroundPipelineScrubber.java:stop(140)) - Stopping Pipeline Scrubber 
Service.
2022-04-27 17:20:47,905 [PipelineScrubberThread] WARN  
pipeline.BackgroundPipelineScrubber (BackgroundPipelineScrubber.java:run(158)) 
- PipelineScrubberThread is interrupted, exit


Wanted but not invoked:
pipelineManager.scrubPipelines();
-> at 
org.apache.hadoop.hdds.scm.pipeline.TestBackgroundPipelineScrubber.testRun(TestBackgroundPipelineScrubber.java:97)
Actually, there were zero interactions with this mock.

Wanted but not invoked:
pipelineManager.scrubPipelines();
-> at 
org.apache.hadoop.hdds.scm.pipeline.TestBackgroundPipelineScrubber.testRun(TestBackgroundPipelineScrubber.java:97)
Actually, there were zero interactions with this mock.
{code}

I believe the reason is that when the `notifyStatusChanged()` method is called, 
the thread may be running and not waiting.

Then calling notifyAll() does not do anything. The thread will fall into the 
wait and stay stuck there until the wait interval expires, or another notify() 
call is received.

After the safemode interval expires, and we received the notifyStatusChanged() 
call in BackgroundPipelineScrubber - do we want the processing thread to wake 
up and run immediately? If not, it could potentially sleep for its wait 
interval after the event is received, which means the delay is actually 
safemode_interval + thread_wait_interval/

I think we need another volatile boolean in `notifyStatusChanged()`, called 
runImmediately. Set it to true and then call notify:

{code}
  synchronized(this) {
    runImmediately = true;
    notify();
  }
{code}

Then in the run loop:

{code}
  synchronized (this) {
    if (!runImmediately) {
      wait(intervalInMillis);
    }
    runImmediately = false;
  }
{code}

This should handle the case where the thread waits just before the notify and I 
believe will fix the test too.




--
This message was sent by Atlassian Jira
(v8.20.7#820007)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to