[ 
https://issues.apache.org/jira/browse/PHOENIX-7980?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Andrew Kyle Purtell updated PHOENIX-7980:
-----------------------------------------
    Description: 
Several Phoenix test classes run with reuseForks=true, sharing a single in-JVM 
minicluster across all classes that land on that same fork (via 
BaseTest.setUpTestDriver). This is important for reducing overall test runtime 
but can be problematic. When the fork's in-JVM minicluster dies this poisons 
every subsequent class assigned to that same reused fork and causes all manner 
of malfunction that currently are not handled all that well.

We should at least bound the failure blast radius with the following basic test 
improvements:

- Add {{<forkedProcessTimeoutInSeconds>1800</forkedProcessTimeoutInSeconds>}} 
to appropriate surefire/failsafe configuration. 
- Add a global ClassRule {{Timeout(20, MINUTES)}} to {{BaseTest}}. Triggers 
from inside the fork and produces a failure report XML for the wedged test 
class.
- Detect and fail fast on discovery of a dead in-JVM cluster in a shared fork. 
In {{BaseTest.doSetup}} after the utility starts, add a health check that 
verifies that the master is up. This alone may reduce failure case wastage of 
test runner resources to seconds from hours.

  was:
Several Phoenix test classes run with reuseForks=true, sharing a single in-JVM 
minicluster across all classes that land on that same fork (via 
BaseTest.setUpTestDriver). This is important for reducing overall test runtime 
but can be problematic. When the fork's in-JVM minicluster dies this poisons 
every subsequent class assigned to that same reused fork and causes all manner 
of malfunction that currently are not handled all that well.

We should at least bound the failure blast radius with the following basic test 
improvements:

- Add {{<forkedProcessTimeoutInSeconds>1800</forkedProcessTimeoutInSeconds>}} 
to surefire/failsafe configuration. This kills any fork that hasn't finished 
its class in 30 min. 
- Add a global ClassRule {{Timeout(20, MINUTES)}} to {{BaseTest}}. Triggers 
from inside the fork and produces a failure report XML for the wedged test 
class.
- Detect and fail fast on discovery of a dead in-JVM cluster in a shared fork. 
In {{BaseTest.doSetup}} after the utility starts, add a health check that 
verifies that the master is up. This alone may reduce failure case wastage of 
test runner resources to seconds from hours.


> Improve handling of test timeouts and resource loss
> ---------------------------------------------------
>
>                 Key: PHOENIX-7980
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-7980
>             Project: Phoenix
>          Issue Type: Bug
>          Components: test
>            Reporter: Andrew Kyle Purtell
>            Assignee: Andrew Kyle Purtell
>            Priority: Major
>
> Several Phoenix test classes run with reuseForks=true, sharing a single 
> in-JVM minicluster across all classes that land on that same fork (via 
> BaseTest.setUpTestDriver). This is important for reducing overall test 
> runtime but can be problematic. When the fork's in-JVM minicluster dies this 
> poisons every subsequent class assigned to that same reused fork and causes 
> all manner of malfunction that currently are not handled all that well.
> We should at least bound the failure blast radius with the following basic 
> test improvements:
> - Add {{<forkedProcessTimeoutInSeconds>1800</forkedProcessTimeoutInSeconds>}} 
> to appropriate surefire/failsafe configuration. 
> - Add a global ClassRule {{Timeout(20, MINUTES)}} to {{BaseTest}}. Triggers 
> from inside the fork and produces a failure report XML for the wedged test 
> class.
> - Detect and fail fast on discovery of a dead in-JVM cluster in a shared 
> fork. In {{BaseTest.doSetup}} after the utility starts, add a health check 
> that verifies that the master is up. This alone may reduce failure case 
> wastage of test runner resources to seconds from hours.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to