[
https://issues.apache.org/jira/browse/PHOENIX-8010?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Xavier Fernandis updated PHOENIX-8010:
--------------------------------------
Description:
h2. Root cause
The crash is in {{getOffsetScanner(...)}} in:
{{phoenix-core-server/src/main/java/org/apache/phoenix/iterate/NonAggregateRegionScannerFactory.java}}
{{}}
{code:java}
byte[] prevScanStartRowKey =
scan.getAttribute(BaseScannerRegionObserverConstants.SCAN_ACTUAL_START_ROW);
// If the region has moved after server has returned dummy or valid row to
client,
// prevScanStartRowKey would be different from actual scan start rowkey.
// ...
if (
Bytes.compareTo(prevScanStartRowKey, initStartRowKey) != 0 && Bytes
.compareTo(ByteUtil.concat(prevScanStartRowKey, ByteUtil.ZERO_BYTE),
initStartRowKey) != 0
) {
iterator.setRowCountToOffset();
}{code}
{{ }}
{{prevScanStartRowKey}} is the value of the {{SCAN_ACTUAL_START_ROW}} . On the
first RPC of a point-lookup query this attribute was never set on the client,
so it is {{{}null{}}}.
* {{org.apache.hadoop.hbase.util.Bytes.compareTo(byte[], byte[])}} is
null-tolerant, so the first comparison does *not* throw.
* {{org.apache.phoenix.util.ByteUtil.concat(byte[] first, byte[]... rest)}} is
*not* null-tolerant. It dereferences {{first.length}} immediately, so
{{concat(null, ByteUtil.ZERO_BYTE)}} throws.
Offending line in:
{{phoenix-core-client/src/main/java/org/apache/phoenix/util/ByteUtil.java}}
{code:java}
public static byte[] concat(byte[] first, byte[]... rest) {
int totalLength = first.length; // <-- NPE when first == null ...
}
{code}
Every range / salted / local-index / guide-post-split scan is produced through
{{{}intersectScan(...){}}}, so those paths always carry the attribute. But the
*point-lookup fast path* in {{{}BaseResultIterators.getParallelScans(byte[],
byte[]){}}}:
{code:java}
if (!isLocalIndex && scanRanges.isPointLookup() &&
!scanRanges.useSkipScanFilter()) {
// builds the scan directly from context.getScan(); never calls
intersectScan(...)
// ... and therefore never sets SCAN_ACTUAL_START_ROW
}
{code}
{{ }}
was:
h2. Root cause
The crash is in {{getOffsetScanner(...)}} in:
{{phoenix-core-server/src/main/java/org/apache/phoenix/iterate/NonAggregateRegionScannerFactory.java}}
{{}}
{code:java}
byte[] prevScanStartRowKey =
scan.getAttribute(BaseScannerRegionObserverConstants.SCAN_ACTUAL_START_ROW);
// If the region has moved after server has returned dummy or valid row to
client,
// prevScanStartRowKey would be different from actual scan start rowkey.
// ...
if (
Bytes.compareTo(prevScanStartRowKey, initStartRowKey) != 0 && Bytes
.compareTo(ByteUtil.concat(prevScanStartRowKey, ByteUtil.ZERO_BYTE),
initStartRowKey) != 0
) {
iterator.setRowCountToOffset();
}{code}
{{ }}
{{prevScanStartRowKey}} is the value of the {{SCAN_ACTUAL_START_ROW}} . On the
first RPC of a point-lookup query this attribute was never set on the client,
so it is {{{}null{}}}.
* {{org.apache.hadoop.hbase.util.Bytes.compareTo(byte[], byte[])}} is
null-tolerant, so the first comparison does *not* throw.
* {{org.apache.phoenix.util.ByteUtil.concat(byte[] first, byte[]... rest)}} is
*not* null-tolerant. It dereferences {{first.length}} immediately, so
{{concat(null, ByteUtil.ZERO_BYTE)}} throws.
Offending line in:
{{phoenix-core-client/src/main/java/org/apache/phoenix/util/ByteUtil.java}}
{code:java}
public static byte[] concat(byte[] first, byte[]... rest) {
int totalLength = first.length; // <-- NPE when first == null ...
}
{code}
Every range / salted / local-index / guide-post-split scan is produced through
{{{}intersectScan(...){}}}, so those paths always carry the attribute. But the
*point-lookup fast path* in {{{}BaseResultIterators.getParallelScans(byte[],
byte[]){}}}:
{{}}
{code:java}
if (!isLocalIndex && scanRanges.isPointLookup() &&
!scanRanges.useSkipScanFilter()) {
// builds the scan directly from context.getScan(); never calls
intersectScan(...)
// ... and therefore never sets SCAN_ACTUAL_START_ROW
}
{code}
{{ }}
{{}}
> NPE in NonAggregateRegionScannerFactory.getOffsetScanner() for point-lookup
> queries with OFFSET
> -----------------------------------------------------------------------------------------------
>
> Key: PHOENIX-8010
> URL: https://issues.apache.org/jira/browse/PHOENIX-8010
> Project: Phoenix
> Issue Type: Bug
> Affects Versions: 5.2.2, 5.3.2
> Environment: h2. Environment
> * Reproduced on Phoenix branch-5.3.0 with HBase 2.6.4.
> * Code path is byte-identical across all 5.x branches and master; HBase
> version is irrelevant.
> Reporter: Xavier Fernandis
> Assignee: Xavier Fernandis
> Priority: Major
>
> h2. Root cause
> The crash is in {{getOffsetScanner(...)}} in:
> {{phoenix-core-server/src/main/java/org/apache/phoenix/iterate/NonAggregateRegionScannerFactory.java}}
> {{}}
> {code:java}
> byte[] prevScanStartRowKey =
> scan.getAttribute(BaseScannerRegionObserverConstants.SCAN_ACTUAL_START_ROW);
> // If the region has moved after server has returned dummy or valid row to
> client,
> // prevScanStartRowKey would be different from actual scan start rowkey.
> // ...
> if (
> Bytes.compareTo(prevScanStartRowKey, initStartRowKey) != 0 && Bytes
> .compareTo(ByteUtil.concat(prevScanStartRowKey, ByteUtil.ZERO_BYTE),
> initStartRowKey) != 0
> ) {
> iterator.setRowCountToOffset();
> }{code}
> {{ }}
> {{prevScanStartRowKey}} is the value of the {{SCAN_ACTUAL_START_ROW}} . On
> the first RPC of a point-lookup query this attribute was never set on the
> client, so it is {{{}null{}}}.
> * {{org.apache.hadoop.hbase.util.Bytes.compareTo(byte[], byte[])}} is
> null-tolerant, so the first comparison does *not* throw.
> * {{org.apache.phoenix.util.ByteUtil.concat(byte[] first, byte[]... rest)}}
> is *not* null-tolerant. It dereferences {{first.length}} immediately, so
> {{concat(null, ByteUtil.ZERO_BYTE)}} throws.
> Offending line in:
> {{phoenix-core-client/src/main/java/org/apache/phoenix/util/ByteUtil.java}}
>
> {code:java}
>
> public static byte[] concat(byte[] first, byte[]... rest) {
> int totalLength = first.length; // <-- NPE when first == null ...
> }
> {code}
>
> Every range / salted / local-index / guide-post-split scan is produced
> through {{{}intersectScan(...){}}}, so those paths always carry the
> attribute. But the *point-lookup fast path* in
> {{{}BaseResultIterators.getParallelScans(byte[], byte[]){}}}:
> {code:java}
> if (!isLocalIndex && scanRanges.isPointLookup() &&
> !scanRanges.useSkipScanFilter()) {
> // builds the scan directly from context.getScan(); never calls
> intersectScan(...)
> // ... and therefore never sets SCAN_ACTUAL_START_ROW
> }
> {code}
> {{ }}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)