[
https://issues.apache.org/jira/browse/SPARK-59379?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
David Mollitor updated SPARK-59379:
-----------------------------------
Description:
h2. Summary
Several {{UTF8String}} methods convert a character index to/from a byte offset
by walking the string one code point at a time ({{{}i +=
numBytesForFirstByte(getByte(i{}}}). For a full-ASCII string every code point
is exactly one byte, so *character index == byte index* and that scan can be
replaced with O(1) arithmetic.
This adds an ASCII fast-path to the four char<->byte "locate" methods:
{{{}substring(int, int){}}}, {{{}getChar(int){}}}, {{{}charPosToByte(int){}}},
and {{{}bytePosToChar(int){}}}.
h2. Details
The fast-path is gated on the cached ASCII flag being {*}already warm{*}, read
without forcing a scan:
{code:java}
if (isFullAscii == IsFullAscii.FULL_ASCII) {
// ASCII: char index == byte index, so locate directly.
...
}
{code}
* This is a pure win with {*}no regression{*}: when the flag is cold
({{{}UNKNOWN{}}}) the methods fall back to the existing per-code-point scan. It
deliberately does NOT call {{{}isFullAscii(){}}}, which would force a full scan
and could regress cold long strings (e.g. {{{}SUBSTR(s, 1, 5){}}}).
* Arithmetic replicates the scan's behavior exactly: a negative index clamps
to 0, the {{until == Integer.MAX_VALUE}} sentinel maps to {{{}numBytes{}}}, and
endpoints clamp to {{{}numBytes{}}}. For example {{substring}} computes {{j =
max(start, 0)}} and {{{}i = (until == Integer.MAX_VALUE) ? numBytes :
min(until, numBytes){}}}; {{getChar}} returns {{codePointFrom(charIndex)}}
directly (the existing {{Objects.checkIndex}} guarantees the index is in range).
* Non-ASCII and invalid-UTF-8 strings have the flag set to {{NOT_ASCII}} when
warmed, so they take the unchanged scan path.
was:
h2. Summary
Several {{UTF8String}} methods convert a character index to/from a byte offset
by walking the string one code point at a time ({{{}i +=
numBytesForFirstByte(getByte(i){}}}). For a full-ASCII string every code point
is exactly one byte, so *character index == byte index* and that scan can be
replaced with O(1) arithmetic.
This adds an ASCII fast-path to the four char<->byte "locate" methods:
{{{}substring(int, int){}}}, {{{}getChar(int){}}}, {{{}charPosToByte(int){}}},
and {{{}bytePosToChar(int){}}}.
h2. Details
The fast-path is gated on the cached ASCII flag being {*}already warm{*}, read
without forcing a scan:
{code:java}
if (isFullAscii == IsFullAscii.FULL_ASCII) {
// ASCII: char index == byte index, so locate directly.
...
}
{code}
* This is a pure win with {*}no regression{*}: when the flag is cold
({{{}UNKNOWN{}}}) the methods fall back to the existing per-code-point scan. It
deliberately does NOT call {{{}isFullAscii(){}}}, which would force a full scan
and could regress cold long strings (e.g. {{{}SUBSTR(s, 1, 5){}}}).
* Arithmetic replicates the scan's behavior exactly: a negative index clamps
to 0, the {{until == Integer.MAX_VALUE}} sentinel maps to {{{}numBytes{}}}, and
endpoints clamp to {{{}numBytes{}}}. For example {{substring}} computes {{j =
max(start, 0)}} and {{{}i = (until == Integer.MAX_VALUE) ? numBytes :
min(until, numBytes){}}}; {{getChar}} returns {{codePointFrom(charIndex)}}
directly (the existing {{Objects.checkIndex}} guarantees the index is in range).
* Non-ASCII and invalid-UTF-8 strings have the flag set to {{NOT_ASCII}} when
warmed, so they take the unchanged scan path.
> Add ASCII fast-path to UTF8String character/byte position lookups
> -----------------------------------------------------------------
>
> Key: SPARK-59379
> URL: https://issues.apache.org/jira/browse/SPARK-59379
> Project: Spark
> Issue Type: Improvement
> Components: Spark Core
> Affects Versions: 4.1.0
> Reporter: David Mollitor
> Priority: Minor
>
> h2. Summary
> Several {{UTF8String}} methods convert a character index to/from a byte
> offset by walking the string one code point at a time ({{{}i +=
> numBytesForFirstByte(getByte(i{}}}). For a full-ASCII string every code point
> is exactly one byte, so *character index == byte index* and that scan can be
> replaced with O(1) arithmetic.
> This adds an ASCII fast-path to the four char<->byte "locate" methods:
> {{{}substring(int, int){}}}, {{{}getChar(int){}}},
> {{{}charPosToByte(int){}}}, and {{{}bytePosToChar(int){}}}.
> h2. Details
> The fast-path is gated on the cached ASCII flag being {*}already warm{*},
> read without forcing a scan:
> {code:java}
> if (isFullAscii == IsFullAscii.FULL_ASCII) {
> // ASCII: char index == byte index, so locate directly.
> ...
> }
> {code}
> * This is a pure win with {*}no regression{*}: when the flag is cold
> ({{{}UNKNOWN{}}}) the methods fall back to the existing per-code-point scan.
> It deliberately does NOT call {{{}isFullAscii(){}}}, which would force a full
> scan and could regress cold long strings (e.g. {{{}SUBSTR(s, 1, 5){}}}).
> * Arithmetic replicates the scan's behavior exactly: a negative index clamps
> to 0, the {{until == Integer.MAX_VALUE}} sentinel maps to {{{}numBytes{}}},
> and endpoints clamp to {{{}numBytes{}}}. For example {{substring}} computes
> {{j = max(start, 0)}} and {{{}i = (until == Integer.MAX_VALUE) ? numBytes :
> min(until, numBytes){}}}; {{getChar}} returns {{codePointFrom(charIndex)}}
> directly (the existing {{Objects.checkIndex}} guarantees the index is in
> range).
> * Non-ASCII and invalid-UTF-8 strings have the flag set to {{NOT_ASCII}}
> when warmed, so they take the unchanged scan path.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]