[
https://issues.apache.org/jira/browse/SOLR-9438?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Shalin Shekhar Mangar updated SOLR-9438:
----------------------------------------
Attachment: SOLR-9438.patch
Patch with a simplified testSplitWithChaosMonkey() that doesn't restart
repeatedly (to avoid the port change problems) and yet reproduces all bugs on
beasting.
I also added another test called testSplitStaticIndexReplication which tickles
the bug related to commit data not being present. A fix for this is also
included.
There are still a few nocommits.
> Shard split can lose data
> -------------------------
>
> Key: SOLR-9438
> URL: https://issues.apache.org/jira/browse/SOLR-9438
> Project: Solr
> Issue Type: Bug
> Security Level: Public(Default Security Level. Issues are Public)
> Components: SolrCloud
> Affects Versions: 4.10.4, 5.5.2, 6.1
> Reporter: Shalin Shekhar Mangar
> Assignee: Shalin Shekhar Mangar
> Labels: difficulty-medium, impact-high
> Fix For: master (7.0), 6.3
>
> Attachments: SOLR-9438-false-replication.log,
> SOLR-9438-split-data-loss.log, SOLR-9438.patch, SOLR-9438.patch,
> SOLR-9438.patch
>
>
> Solr’s shard split can lose documents if the parent/sub-shard leader is
> killed (or crashes) between the time that the new sub-shard replica is
> created and before it recovers. In such a case the slice has already been set
> to ‘recovery’ state, the sub-shard replica comes up, finds that no other
> replica is up, waits until the leader vote wait time and then proceeds to
> become the leader as well as publish itself as active. Once that happens the
> overseer seeing that all replicas of the sub-shard are now ‘active’, sets the
> parent slice as ‘inactive’ and the new sub-shard as ‘active’.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]