[ 
https://issues.apache.org/jira/browse/PHOENIX-8016?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Tanuj Khurana updated PHOENIX-8016:
-----------------------------------
    Description: 
Orphaned/stale uncovered-index entries (e.g. a NULL-keyed entry left behind 
after a row's indexed value changes from NULL to non-null) are supposed to be 
cleaned by lazy read-repair, but for an actively-updated data row they are 
*never* deleted, no matter how old the stale index cell is.

Read-repair for uncovered indexes runs on the read path in
{{UncoveredIndexRegionScanner.verifyIndexRowAndRepairIfNecessary}} (:321-362) 
and
the {{dataRow == null}} branch of {{getNextCoveredIndexRow}} (:384-393), gated 
on
{{ageThreshold}} ({{phoenix.global.index.row.age.threshold.to.delete.ms}},
default 7 days, {{QueryServicesOptions.java:439-440}}).

A stale (but non-orphan) entry hits the "data row exists but index row is stale"
case (:353-360). The tenant's data row still exists (now with a non-null indexed
value), so {{checkIndexRow}} fails, and the delete is gated on *the data row's*
max timestamp:

{code:java}
if (indexMaintainer.isAgedEnough(IndexUtil.getMaxTimestamp(put), ageThreshold) 
...) {   // :354
  region.delete(indexMaintainer.createDelete(indexRowKey, 
IndexUtil.getMaxTimestamp(put), false));

If the data table keeps getting updated frequently, the isAgedEnough check will 
never be true because it takes the latest timestamp of the data table row 
instead of the orphaned index row

  was:
Orphaned/stale uncovered-index entries (e.g. a NULL-keyed entry left behind 
after
a row's indexed value changes from NULL to non-null) are supposed to be cleaned 
by
lazy read-repair, but for an actively-updated data row they are *never* deleted,
no matter how old the stale index cell is.

Read-repair for uncovered indexes runs on the read path in
{{UncoveredIndexRegionScanner.verifyIndexRowAndRepairIfNecessary}} (:321-362) 
and
the {{dataRow == null}} branch of {{getNextCoveredIndexRow}} (:384-393), gated 
on
{{ageThreshold}} ({{phoenix.global.index.row.age.threshold.to.delete.ms}},
default 7 days, {{QueryServicesOptions.java:439-440}}).

A stale (but non-orphan) entry hits the "data row exists but index row is stale"
case (:353-360). The tenant's data row still exists (now with a non-null indexed
value), so {{checkIndexRow}} fails, and the delete is gated on *the data row's*
max timestamp:

{code:java}
if (indexMaintainer.isAgedEnough(IndexUtil.getMaxTimestamp(put), ageThreshold) 
...) {   // :354
  region.delete(indexMaintainer.createDelete(indexRowKey, 
IndexUtil.getMaxTimestamp(put), false));


> Uncovered-index read-repair never deletes a stale index entry when the data 
> row is actively updated
> ---------------------------------------------------------------------------------------------------
>
>                 Key: PHOENIX-8016
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-8016
>             Project: Phoenix
>          Issue Type: Bug
>    Affects Versions: 5.2.0, 5.2.1, 5.3.0, 5.2.2, 5.3.1, 5.3.2
>            Reporter: Tanuj Khurana
>            Priority: Major
>
> Orphaned/stale uncovered-index entries (e.g. a NULL-keyed entry left behind 
> after a row's indexed value changes from NULL to non-null) are supposed to be 
> cleaned by lazy read-repair, but for an actively-updated data row they are 
> *never* deleted, no matter how old the stale index cell is.
> Read-repair for uncovered indexes runs on the read path in
> {{UncoveredIndexRegionScanner.verifyIndexRowAndRepairIfNecessary}} (:321-362) 
> and
> the {{dataRow == null}} branch of {{getNextCoveredIndexRow}} (:384-393), 
> gated on
> {{ageThreshold}} ({{phoenix.global.index.row.age.threshold.to.delete.ms}},
> default 7 days, {{QueryServicesOptions.java:439-440}}).
> A stale (but non-orphan) entry hits the "data row exists but index row is 
> stale"
> case (:353-360). The tenant's data row still exists (now with a non-null 
> indexed
> value), so {{checkIndexRow}} fails, and the delete is gated on *the data 
> row's*
> max timestamp:
> {code:java}
> if (indexMaintainer.isAgedEnough(IndexUtil.getMaxTimestamp(put), 
> ageThreshold) ...) {   // :354
>   region.delete(indexMaintainer.createDelete(indexRowKey, 
> IndexUtil.getMaxTimestamp(put), false));
> If the data table keeps getting updated frequently, the isAgedEnough check 
> will never be true because it takes the latest timestamp of the data table 
> row instead of the orphaned index row



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to