[
https://issues.apache.org/jira/browse/HBASE-18152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
stack updated HBASE-18152:
--------------------------
Attachment: HBASE-17537.master.002.patch
Patch was too stringent. Was reporting the below as NOT increasing when it is:
{code}
2017-06-02 17:17:53,113 WARN [ve0524:16000.masterManager]
wal.ProcedureWALFormatReader: NOT INCREASING! current=class_name:
"org.apache.hadoop.hbase.master.procedure.ServerCrashProcedure"
proc_id: 21
submitted_time: 1496449031443
owner: "stack"
state: WAITING
stack_id: 0
stack_id: 1
stack_id: 2
stack_id: 3
last_update: 1496449035902
state_data:
"\n\b\001\b\003\b\005\b\b\bdR\n%\n\031ve0524.halxg.cloudera.com\020\200}\030\344\371\256\332\306+\032%\b\215\341\225\331\306+\022\022\n\005hbase\022\tnamespace\032\000\"\000(\0000\0008\000(\0000\001"
, candidate=class_name:
"org.apache.hadoop.hbase.master.procedure.ServerCrashProcedure"
proc_id: 21
submitted_time: 1496449031443
owner: "stack"
state: RUNNABLE
stack_id: 0
stack_id: 1
stack_id: 2
stack_id: 3
last_update: 1496449036513
state_data:
"\n\b\001\b\003\b\005\b\b\bdR\n%\n\031ve0524.halxg.cloudera.com\020\200}\030\344\371\256\332\306+\032%\b\215\341\225\331\306+\022\022\n\005hbase\022\tnamespace\032\000\"\000(\0000\0008\000(\0000\001"
{code}
... notice the timestamps.
> [AMv2] Corrupt Procedure WAL file; procedure data stored out of order
> ---------------------------------------------------------------------
>
> Key: HBASE-18152
> URL: https://issues.apache.org/jira/browse/HBASE-18152
> Project: HBase
> Issue Type: Bug
> Components: Region Assignment
> Affects Versions: 2.0.0
> Reporter: stack
> Assignee: stack
> Priority: Critical
> Fix For: 2.0.0
>
> Attachments: HBASE-17537.master.002.patch,
> HBASE-18152.master.001.patch, pv2-00000000000000000047.log,
> reading_bad_wal.patch
>
>
> I've seen corruption from time-to-time testing. Its rare enough. Often we
> can get over it but sometimes we can't. It took me a while to capture an
> instance of corruption. Turns out we are write to the WAL out-of-order which
> undoes a basic tenet; that WAL content is ordered in line w/ execution.
> Below I'll post a corrupt WAL.
> Looking at the write-side, there is a lot going on. I'm not clear on how we
> could write out of order. Will try and get more insight. Meantime parking
> this issue here to fill data into.
--
This message was sent by Atlassian JIRA
(v6.3.15#6346)