[ 
https://issues.apache.org/jira/browse/HBASE-18152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

stack updated HBASE-18152:
--------------------------
    Attachment: HBASE-17537.master.002.patch

Patch was too stringent. Was reporting the below as NOT increasing when it is:
{code}
2017-06-02 17:17:53,113 WARN  [ve0524:16000.masterManager] 
wal.ProcedureWALFormatReader: NOT INCREASING! current=class_name: 
"org.apache.hadoop.hbase.master.procedure.ServerCrashProcedure"
proc_id: 21
submitted_time: 1496449031443
owner: "stack"
state: WAITING
stack_id: 0
stack_id: 1
stack_id: 2
stack_id: 3
last_update: 1496449035902
state_data: 
"\n\b\001\b\003\b\005\b\b\bdR\n%\n\031ve0524.halxg.cloudera.com\020\200}\030\344\371\256\332\306+\032%\b\215\341\225\331\306+\022\022\n\005hbase\022\tnamespace\032\000\"\000(\0000\0008\000(\0000\001"
, candidate=class_name: 
"org.apache.hadoop.hbase.master.procedure.ServerCrashProcedure"
proc_id: 21
submitted_time: 1496449031443
owner: "stack"
state: RUNNABLE
stack_id: 0
stack_id: 1
stack_id: 2
stack_id: 3
last_update: 1496449036513
state_data: 
"\n\b\001\b\003\b\005\b\b\bdR\n%\n\031ve0524.halxg.cloudera.com\020\200}\030\344\371\256\332\306+\032%\b\215\341\225\331\306+\022\022\n\005hbase\022\tnamespace\032\000\"\000(\0000\0008\000(\0000\001"
{code}

... notice the timestamps.

> [AMv2] Corrupt Procedure WAL file; procedure data stored out of order
> ---------------------------------------------------------------------
>
>                 Key: HBASE-18152
>                 URL: https://issues.apache.org/jira/browse/HBASE-18152
>             Project: HBase
>          Issue Type: Bug
>          Components: Region Assignment
>    Affects Versions: 2.0.0
>            Reporter: stack
>            Assignee: stack
>            Priority: Critical
>             Fix For: 2.0.0
>
>         Attachments: HBASE-17537.master.002.patch, 
> HBASE-18152.master.001.patch, pv2-00000000000000000047.log, 
> reading_bad_wal.patch
>
>
> I've seen corruption from time-to-time testing.  Its rare enough. Often we 
> can get over it but sometimes we can't. It took me a while to capture an 
> instance of corruption. Turns out we are write to the WAL out-of-order which 
> undoes a basic tenet; that WAL content is ordered in line w/ execution.
> Below I'll post a corrupt WAL.
> Looking at the write-side, there is a lot going on. I'm not clear on how we 
> could write out of order. Will try and get more insight. Meantime parking 
> this issue here to fill data into.



--
This message was sent by Atlassian JIRA
(v6.3.15#6346)

Reply via email to