[
https://issues.apache.org/jira/browse/HBASE-21463?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16682845#comment-16682845
]
Duo Zhang commented on HBASE-21463:
-----------------------------------
OK the problem is that, we skip the table state check and schedule TRSP to
assign a region when it is already online, then we stuck there as when the
region server received the OpenRegion request it finds out that the region is
already online, so just ignores. In the old time the regionServerReport can
help us recovering from the stuck, but in the patch here we removed this
behavior.
This does not make sense. We should always check the table state. If there are
broken states or some other errors, we should use HBCK2 to fix them.
So I plan to remove the usage of this flag, and also remove the UT.
> The checkOnlineRegionsReport can accidentally complete a TRSP
> -------------------------------------------------------------
>
> Key: HBASE-21463
> URL: https://issues.apache.org/jira/browse/HBASE-21463
> Project: HBase
> Issue Type: Sub-task
> Components: amv2
> Reporter: Duo Zhang
> Assignee: Duo Zhang
> Priority: Critical
> Fix For: 3.0.0, 2.2.0
>
> Attachments: HBASE-21463-UT.patch, HBASE-21463.patch
>
>
> On our testing cluster, we observe a race condition:
> 1. A regionServerReport request is built
> 2. A TRSP is scheduled to reopen the region
> 3. The region is closed at RS side
> 4. The OpenRegionProcedure is created
> 5. The regionServerReport generated at step 1 is executed, and we find that
> the region is opened on the RS, so we update the region state to OPEN.
> 6. The OpenRegionProcedure notices that the region has already been in the
> OPEN state so gives up and finishes.
> 7. The TRSP finishes.
> 8. The region is recorded as OPEN on the RS but actually not, and can not
> recover unless we use HBCK2.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)