Nirdosh Kumar Yadav created HBASE-30335:
-------------------------------------------
Summary: Master's flushedSequenceIdByRegion is not seeded with
openSeqNum on region OPEN, allowing WAL split to write stale recovered.edits
for edits already durable in HFiles
Key: HBASE-30335
URL: https://issues.apache.org/jira/browse/HBASE-30335
Project: HBase
Issue Type: Bug
Components: master
Affects Versions: 2.5.10, 2.6.1
Reporter: Nirdosh Kumar Yadav
Assignee: Nirdosh Kumar Yadav
h3. Problem Statement
When a region is drain-moved from RS-A to RS-B and RS-A crashes shortly
afterwards while its WALs still contain (already-durable) edits for that
region, the subsequent ServerCrashProcedure WAL split writes a stale
recovered.edits/<seqId> file under RS-B's already-online region directory. The
file contains edits that are already in HFiles, but downstream guards treat any
recovered.edits/<numeric> file as an unreplayed durability risk. The next
MergeTableRegionsProcedure (or SplitTableRegionProcedure) on that region trips
AssignmentManagerUtil.checkClosedRegion() with "Recovered.edits are found in
Region: … abort split/merge to prevent data loss" and leaves the region stuck
in RIT until an operator intervenes.
h3. Root Cause
WALSplitter decides whether to write an edit to recovered.edits/ by comparing
the edit's sequenceId against the region's lastFlushedSequenceId returned by
the master (WALSplitter.java:358-360). The master's
ServerManager.flushedSequenceIdByRegion map is populated only by RegionServer
heartbeats (ServerManager.java:207-223). During a drain-move:
* The source RS drops the region from onlineRegions after CLOSE and stops
heartbeating for it.
* The destination RS has just OPENed the region and has no new writes to
flush, so it has nothing to report.
* No code path refreshes the master's cache from the OPEN transition.
Consequently, when WALSplitter later asks the master for the region's
lastFlushedSequenceId, it gets NO_SEQNUM and falls back to -1L at
WALSplitter.java:358-360:
{code:java}
if (lastFlushedSequenceId == null) { lastFlushedSequenceId = -1L; // Treat as
"never flushed" — write everything. } {code}
Every edit in the crashed server's WAL for that region is now above the
fallback watermark and gets written to recovered.edits/, even though the
region's on-disk .seqid marker already proves those edits are durable.
The single master-side transition point to State.OPEN is
AssignmentManager.regionOpenedWithoutPersistingToMeta()
(AssignmentManager.java:2150), where regionNode.getOpenSeqNum() is already
available from the RS's OPEN transition report. That value is greater than
every durable seqId at the previous CLOSE and is exactly the watermark
WALSplitter needs — but it is never propagated to
ServerManager.flushedSequenceIdByRegion.
h3. Key Observations
* Region 112d9f08181e80aaae01e68c9229c3cb was gracefully drained from rs-132 →
rs-81 at 16:36:34.001 UTC; rs-81's openSeqNum=4997750282. The close on rs-132
had already flushed all edits.
* ~25 seconds later rs-132 was declared dead as part of a broader RS drain.
* WAL-split worker rs-34 wrote recovered.edits/0000000004997750280 at
16:37:12.712 UTC — 38.7 seconds after rs-81 was live on the region. The edit's
seqId=4997750280 was already durable in HFiles (durable watermark ≥ 4997750281).
* The orphan sat dormant for 1h 55m until MergeTableRegionsProcedure
pid=46990253 tripped checkClosedRegion at 18:32:33.293 UTC and the region
stayed in RIT for 48m 56s.
* When rs-120 eventually replayed the file after operator intervention, HBase
logged Applied 0, skipped 1 — direct proof that the edit was already durable
and the guard was a false-positive.
h3. Proposed Solution
Seed the master's flushed-sequence cache from the OPEN transition. In
AssignmentManager.java:2150:
{code:java}
void regionOpenedWithoutPersistingToMeta(RegionStateNode regionNode) throws
IOException { regionNode.transitionState(State.OPEN, STATES_EXPECTED_ON_OPEN);
RegionInfo regionInfo = regionNode.getRegionInfo();
regionStates.addRegionToServer(regionNode);
regionStates.removeFromFailedOpen(regionInfo); +
master.getServerManager().reportRegionOpen(regionInfo,
regionNode.getOpenSeqNum()); } {code}
Companion in ServerManager.java (adjacent to updateLastFlushedSequenceIds at
line 207):
{code:java}
public void reportRegionOpen(RegionInfo regionInfo, long openSeqNum) { if
(openSeqNum == HConstants.NO_SEQNUM) return; byte[] encodedRegionName =
Bytes.toBytes(regionInfo.getEncodedName()); // openSeqNum > every durable seqId
at CLOSE time — preserves the // "keep the max" invariant that
updateLastFlushedSequenceIds already uses.
flushedSequenceIdByRegion.merge(encodedRegionName, openSeqNum, Math::max); }
{code}
Effect. At OPEN time the master learns the correct durable watermark
synchronously with the state transition. When WALSplitter later queries
ids.getLastFlushedSequenceId(), it gets openSeqNum instead of NO_SEQNUM; the
skip check openSeqNum >= edit.seqId evaluates true for every already-durable
edit, and no recovered.edits/* file is written. The 48-min RIT class of failure
is eliminated at its source — no relaxation of the downstream recovered.edits
guard is required.
{code:java}
if (lastFlushedSequenceId == null || lastFlushedSequenceId ==
HConstants.NO_SEQNUM) { Path regionDir = FSUtils.getRegionDirFromTableDir(
CommonFSUtils.getTableDir(walRootDir, entry.getKey().getTableName()),
encodedRegionNameAsStr); long onDiskMaxSeqId =
WALSplitUtil.getMaxRegionSequenceId(walFS, regionDir); lastFlushedSequenceId =
(onDiskMaxSeqId != HConstants.NO_SEQNUM) ? onDiskMaxSeqId : -1L; } {code}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)