[
https://issues.apache.org/jira/browse/PHOENIX-8016?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Tanuj Khurana updated PHOENIX-8016:
-----------------------------------
Description:
Orphaned/stale uncovered-index entries (e.g. a NULL-keyed entry left behind
after a row's indexed value changes from NULL to non-null) are supposed to be
cleaned by lazy read-repair, but for an actively-updated data row they are
*never* deleted, no matter how old the stale index cell is.
Read-repair for uncovered indexes runs on the read path in
{{UncoveredIndexRegionScanner.verifyIndexRowAndRepairIfNecessary}} (:321-362)
and
the {{dataRow == null}} branch of {{getNextCoveredIndexRow}} (:384-393), gated
on
{{ageThreshold}} ({{phoenix.global.index.row.age.threshold.to.delete.ms}},
default 7 days, {{QueryServicesOptions.java:439-440}}).
A stale (but non-orphan) entry hits the "data row exists but index row is stale"
case (:353-360). The tenant's data row still exists (now with a non-null indexed
value), so {{checkIndexRow}} fails, and the delete is gated on *the data row's*
max timestamp:
{code:java}
if (indexMaintainer.isAgedEnough(IndexUtil.getMaxTimestamp(put), ageThreshold)
...) { // :354
region.delete(indexMaintainer.createDelete(indexRowKey,
IndexUtil.getMaxTimestamp(put), false));
If the data table keeps getting updated frequently, the isAgedEnough check will
never be true because it takes the latest timestamp of the data table row
instead of the orphaned index row
was:
Orphaned/stale uncovered-index entries (e.g. a NULL-keyed entry left behind
after
a row's indexed value changes from NULL to non-null) are supposed to be cleaned
by
lazy read-repair, but for an actively-updated data row they are *never* deleted,
no matter how old the stale index cell is.
Read-repair for uncovered indexes runs on the read path in
{{UncoveredIndexRegionScanner.verifyIndexRowAndRepairIfNecessary}} (:321-362)
and
the {{dataRow == null}} branch of {{getNextCoveredIndexRow}} (:384-393), gated
on
{{ageThreshold}} ({{phoenix.global.index.row.age.threshold.to.delete.ms}},
default 7 days, {{QueryServicesOptions.java:439-440}}).
A stale (but non-orphan) entry hits the "data row exists but index row is stale"
case (:353-360). The tenant's data row still exists (now with a non-null indexed
value), so {{checkIndexRow}} fails, and the delete is gated on *the data row's*
max timestamp:
{code:java}
if (indexMaintainer.isAgedEnough(IndexUtil.getMaxTimestamp(put), ageThreshold)
...) { // :354
region.delete(indexMaintainer.createDelete(indexRowKey,
IndexUtil.getMaxTimestamp(put), false));
> Uncovered-index read-repair never deletes a stale index entry when the data
> row is actively updated
> ---------------------------------------------------------------------------------------------------
>
> Key: PHOENIX-8016
> URL: https://issues.apache.org/jira/browse/PHOENIX-8016
> Project: Phoenix
> Issue Type: Bug
> Affects Versions: 5.2.0, 5.2.1, 5.3.0, 5.2.2, 5.3.1, 5.3.2
> Reporter: Tanuj Khurana
> Priority: Major
>
> Orphaned/stale uncovered-index entries (e.g. a NULL-keyed entry left behind
> after a row's indexed value changes from NULL to non-null) are supposed to be
> cleaned by lazy read-repair, but for an actively-updated data row they are
> *never* deleted, no matter how old the stale index cell is.
> Read-repair for uncovered indexes runs on the read path in
> {{UncoveredIndexRegionScanner.verifyIndexRowAndRepairIfNecessary}} (:321-362)
> and
> the {{dataRow == null}} branch of {{getNextCoveredIndexRow}} (:384-393),
> gated on
> {{ageThreshold}} ({{phoenix.global.index.row.age.threshold.to.delete.ms}},
> default 7 days, {{QueryServicesOptions.java:439-440}}).
> A stale (but non-orphan) entry hits the "data row exists but index row is
> stale"
> case (:353-360). The tenant's data row still exists (now with a non-null
> indexed
> value), so {{checkIndexRow}} fails, and the delete is gated on *the data
> row's*
> max timestamp:
> {code:java}
> if (indexMaintainer.isAgedEnough(IndexUtil.getMaxTimestamp(put),
> ageThreshold) ...) { // :354
> region.delete(indexMaintainer.createDelete(indexRowKey,
> IndexUtil.getMaxTimestamp(put), false));
> If the data table keeps getting updated frequently, the isAgedEnough check
> will never be true because it takes the latest timestamp of the data table
> row instead of the orphaned index row
--
This message was sent by Atlassian Jira
(v8.20.10#820010)