[ 
https://issues.apache.org/jira/browse/PHOENIX-8010?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Xavier Fernandis updated PHOENIX-8010:
--------------------------------------
    Description: 
h2. Root cause

The crash is in {{getOffsetScanner(...)}} in:

{{phoenix-core-server/src/main/java/org/apache/phoenix/iterate/NonAggregateRegionScannerFactory.java}}

{{}}
{code:java}
byte[] prevScanStartRowKey =
  scan.getAttribute(BaseScannerRegionObserverConstants.SCAN_ACTUAL_START_ROW);
// If the region has moved after server has returned dummy or valid row to 
client,
// prevScanStartRowKey would be different from actual scan start rowkey.
// ...
if (
  Bytes.compareTo(prevScanStartRowKey, initStartRowKey) != 0 && Bytes
    .compareTo(ByteUtil.concat(prevScanStartRowKey, ByteUtil.ZERO_BYTE), 
initStartRowKey) != 0
) {
  iterator.setRowCountToOffset();
}{code}
{{ }}

{{prevScanStartRowKey}} is the value of the {{SCAN_ACTUAL_START_ROW}} . On the 
first RPC of a point-lookup query this attribute was never set on the client, 
so it is {{{}null{}}}.
 * {{org.apache.hadoop.hbase.util.Bytes.compareTo(byte[], byte[])}} is 
null-tolerant, so the first comparison does *not* throw.
 * {{org.apache.phoenix.util.ByteUtil.concat(byte[] first, byte[]... rest)}} is 
*not* null-tolerant. It dereferences {{first.length}} immediately, so 
{{concat(null, ByteUtil.ZERO_BYTE)}} throws.

Offending line in:

{{phoenix-core-client/src/main/java/org/apache/phoenix/util/ByteUtil.java}}

 
{code:java}
 
public static byte[] concat(byte[] first, byte[]... rest) {
 int totalLength = first.length; // <-- NPE when first == null ... 
}
{code}
 

Every range / salted / local-index / guide-post-split scan is produced through 
{{{}intersectScan(...){}}}, so those paths always carry the attribute. But the 
*point-lookup fast path* in {{{}BaseResultIterators.getParallelScans(byte[], 
byte[]){}}}:
{code:java}
if (!isLocalIndex && scanRanges.isPointLookup() && 
!scanRanges.useSkipScanFilter()) {
    // builds the scan directly from context.getScan(); never calls 
intersectScan(...)
    // ... and therefore never sets SCAN_ACTUAL_START_ROW
}
{code}
{{ }}

  was:
h2. Root cause

The crash is in {{getOffsetScanner(...)}} in:

{{phoenix-core-server/src/main/java/org/apache/phoenix/iterate/NonAggregateRegionScannerFactory.java}}

{{}}
{code:java}
byte[] prevScanStartRowKey =
  scan.getAttribute(BaseScannerRegionObserverConstants.SCAN_ACTUAL_START_ROW);
// If the region has moved after server has returned dummy or valid row to 
client,
// prevScanStartRowKey would be different from actual scan start rowkey.
// ...
if (
  Bytes.compareTo(prevScanStartRowKey, initStartRowKey) != 0 && Bytes
    .compareTo(ByteUtil.concat(prevScanStartRowKey, ByteUtil.ZERO_BYTE), 
initStartRowKey) != 0
) {
  iterator.setRowCountToOffset();
}{code}
{{ }}

{{prevScanStartRowKey}} is the value of the {{SCAN_ACTUAL_START_ROW}} . On the 
first RPC of a point-lookup query this attribute was never set on the client, 
so it is {{{}null{}}}.
 * {{org.apache.hadoop.hbase.util.Bytes.compareTo(byte[], byte[])}} is 
null-tolerant, so the first comparison does *not* throw.
 * {{org.apache.phoenix.util.ByteUtil.concat(byte[] first, byte[]... rest)}} is 
*not* null-tolerant. It dereferences {{first.length}} immediately, so 
{{concat(null, ByteUtil.ZERO_BYTE)}} throws.

Offending line in:

{{phoenix-core-client/src/main/java/org/apache/phoenix/util/ByteUtil.java}}

 
{code:java}
 
public static byte[] concat(byte[] first, byte[]... rest) {
 int totalLength = first.length; // <-- NPE when first == null ... 
}
{code}
 

Every range / salted / local-index / guide-post-split scan is produced through 
{{{}intersectScan(...){}}}, so those paths always carry the attribute. But the 
*point-lookup fast path* in {{{}BaseResultIterators.getParallelScans(byte[], 
byte[]){}}}:

{{}}
{code:java}
if (!isLocalIndex && scanRanges.isPointLookup() && 
!scanRanges.useSkipScanFilter()) {
    // builds the scan directly from context.getScan(); never calls 
intersectScan(...)
    // ... and therefore never sets SCAN_ACTUAL_START_ROW
}
{code}
{{ }}

{{}}


> NPE in NonAggregateRegionScannerFactory.getOffsetScanner() for point-lookup 
> queries with OFFSET
> -----------------------------------------------------------------------------------------------
>
>                 Key: PHOENIX-8010
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-8010
>             Project: Phoenix
>          Issue Type: Bug
>    Affects Versions: 5.2.2, 5.3.2
>         Environment: h2. Environment
>  * Reproduced on Phoenix branch-5.3.0 with HBase 2.6.4.
>  * Code path is byte-identical across all 5.x branches and master; HBase 
> version is irrelevant.
>            Reporter: Xavier Fernandis
>            Assignee: Xavier Fernandis
>            Priority: Major
>
> h2. Root cause
> The crash is in {{getOffsetScanner(...)}} in:
> {{phoenix-core-server/src/main/java/org/apache/phoenix/iterate/NonAggregateRegionScannerFactory.java}}
> {{}}
> {code:java}
> byte[] prevScanStartRowKey =
>   scan.getAttribute(BaseScannerRegionObserverConstants.SCAN_ACTUAL_START_ROW);
> // If the region has moved after server has returned dummy or valid row to 
> client,
> // prevScanStartRowKey would be different from actual scan start rowkey.
> // ...
> if (
>   Bytes.compareTo(prevScanStartRowKey, initStartRowKey) != 0 && Bytes
>     .compareTo(ByteUtil.concat(prevScanStartRowKey, ByteUtil.ZERO_BYTE), 
> initStartRowKey) != 0
> ) {
>   iterator.setRowCountToOffset();
> }{code}
> {{ }}
> {{prevScanStartRowKey}} is the value of the {{SCAN_ACTUAL_START_ROW}} . On 
> the first RPC of a point-lookup query this attribute was never set on the 
> client, so it is {{{}null{}}}.
>  * {{org.apache.hadoop.hbase.util.Bytes.compareTo(byte[], byte[])}} is 
> null-tolerant, so the first comparison does *not* throw.
>  * {{org.apache.phoenix.util.ByteUtil.concat(byte[] first, byte[]... rest)}} 
> is *not* null-tolerant. It dereferences {{first.length}} immediately, so 
> {{concat(null, ByteUtil.ZERO_BYTE)}} throws.
> Offending line in:
> {{phoenix-core-client/src/main/java/org/apache/phoenix/util/ByteUtil.java}}
>  
> {code:java}
>  
> public static byte[] concat(byte[] first, byte[]... rest) {
>  int totalLength = first.length; // <-- NPE when first == null ... 
> }
> {code}
>  
> Every range / salted / local-index / guide-post-split scan is produced 
> through {{{}intersectScan(...){}}}, so those paths always carry the 
> attribute. But the *point-lookup fast path* in 
> {{{}BaseResultIterators.getParallelScans(byte[], byte[]){}}}:
> {code:java}
> if (!isLocalIndex && scanRanges.isPointLookup() && 
> !scanRanges.useSkipScanFilter()) {
>     // builds the scan directly from context.getScan(); never calls 
> intersectScan(...)
>     // ... and therefore never sets SCAN_ACTUAL_START_ROW
> }
> {code}
> {{ }}



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to