[ 
https://issues.apache.org/jira/browse/SPARK-59379?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

David Mollitor updated SPARK-59379:
-----------------------------------
    Description: 
h2. Summary

Several {{UTF8String}} methods convert a character index to/from a byte offset 
by walking the string one code point at a time ({{{}i += 
numBytesForFirstByte(getByte(i{}}}). For a full-ASCII string every code point 
is exactly one byte, so *character index == byte index* and that scan can be 
replaced with O(1) arithmetic.

This adds an ASCII fast-path to the four char<->byte "locate" methods: 
{{{}substring(int, int){}}}, {{{}getChar(int){}}}, {{{}charPosToByte(int){}}}, 
and {{{}bytePosToChar(int){}}}.
h2. Details

The fast-path is gated on the cached ASCII flag being {*}already warm{*}, read 
without forcing a scan:
{code:java}
if (isFullAscii == IsFullAscii.FULL_ASCII) {
  // ASCII: char index == byte index, so locate directly.
  ...
}
{code}
 * This is a pure win with {*}no regression{*}: when the flag is cold 
({{{}UNKNOWN{}}}) the methods fall back to the existing per-code-point scan. It 
deliberately does NOT call {{{}isFullAscii(){}}}, which would force a full scan 
and could regress cold long strings (e.g. {{{}SUBSTR(s, 1, 5){}}}).
 * Arithmetic replicates the scan's behavior exactly: a negative index clamps 
to 0, the {{until == Integer.MAX_VALUE}} sentinel maps to {{{}numBytes{}}}, and 
endpoints clamp to {{{}numBytes{}}}. For example {{substring}} computes {{j = 
max(start, 0)}} and {{{}i = (until == Integer.MAX_VALUE) ? numBytes : 
min(until, numBytes){}}}; {{getChar}} returns {{codePointFrom(charIndex)}} 
directly (the existing {{Objects.checkIndex}} guarantees the index is in range).
 * Non-ASCII and invalid-UTF-8 strings have the flag set to {{NOT_ASCII}} when 
warmed, so they take the unchanged scan path.

  was:
h2. Summary

Several {{UTF8String}} methods convert a character index to/from a byte offset 
by walking the string one code point at a time ({{{}i += 
numBytesForFirstByte(getByte(i){}}}). For a full-ASCII string every code point 
is exactly one byte, so *character index == byte index* and that scan can be 
replaced with O(1) arithmetic.

This adds an ASCII fast-path to the four char<->byte "locate" methods: 
{{{}substring(int, int){}}}, {{{}getChar(int){}}}, {{{}charPosToByte(int){}}}, 
and {{{}bytePosToChar(int){}}}.
h2. Details

The fast-path is gated on the cached ASCII flag being {*}already warm{*}, read 
without forcing a scan:
{code:java}
if (isFullAscii == IsFullAscii.FULL_ASCII) {
  // ASCII: char index == byte index, so locate directly.
  ...
}
{code}
 * This is a pure win with {*}no regression{*}: when the flag is cold 
({{{}UNKNOWN{}}}) the methods fall back to the existing per-code-point scan. It 
deliberately does NOT call {{{}isFullAscii(){}}}, which would force a full scan 
and could regress cold long strings (e.g. {{{}SUBSTR(s, 1, 5){}}}).
 * Arithmetic replicates the scan's behavior exactly: a negative index clamps 
to 0, the {{until == Integer.MAX_VALUE}} sentinel maps to {{{}numBytes{}}}, and 
endpoints clamp to {{{}numBytes{}}}. For example {{substring}} computes {{j = 
max(start, 0)}} and {{{}i = (until == Integer.MAX_VALUE) ? numBytes : 
min(until, numBytes){}}}; {{getChar}} returns {{codePointFrom(charIndex)}} 
directly (the existing {{Objects.checkIndex}} guarantees the index is in range).
 * Non-ASCII and invalid-UTF-8 strings have the flag set to {{NOT_ASCII}} when 
warmed, so they take the unchanged scan path.


> Add ASCII fast-path to UTF8String character/byte position lookups
> -----------------------------------------------------------------
>
>                 Key: SPARK-59379
>                 URL: https://issues.apache.org/jira/browse/SPARK-59379
>             Project: Spark
>          Issue Type: Improvement
>          Components: Spark Core
>    Affects Versions: 4.1.0
>            Reporter: David Mollitor
>            Priority: Minor
>
> h2. Summary
> Several {{UTF8String}} methods convert a character index to/from a byte 
> offset by walking the string one code point at a time ({{{}i += 
> numBytesForFirstByte(getByte(i{}}}). For a full-ASCII string every code point 
> is exactly one byte, so *character index == byte index* and that scan can be 
> replaced with O(1) arithmetic.
> This adds an ASCII fast-path to the four char<->byte "locate" methods: 
> {{{}substring(int, int){}}}, {{{}getChar(int){}}}, 
> {{{}charPosToByte(int){}}}, and {{{}bytePosToChar(int){}}}.
> h2. Details
> The fast-path is gated on the cached ASCII flag being {*}already warm{*}, 
> read without forcing a scan:
> {code:java}
> if (isFullAscii == IsFullAscii.FULL_ASCII) {
>   // ASCII: char index == byte index, so locate directly.
>   ...
> }
> {code}
>  * This is a pure win with {*}no regression{*}: when the flag is cold 
> ({{{}UNKNOWN{}}}) the methods fall back to the existing per-code-point scan. 
> It deliberately does NOT call {{{}isFullAscii(){}}}, which would force a full 
> scan and could regress cold long strings (e.g. {{{}SUBSTR(s, 1, 5){}}}).
>  * Arithmetic replicates the scan's behavior exactly: a negative index clamps 
> to 0, the {{until == Integer.MAX_VALUE}} sentinel maps to {{{}numBytes{}}}, 
> and endpoints clamp to {{{}numBytes{}}}. For example {{substring}} computes 
> {{j = max(start, 0)}} and {{{}i = (until == Integer.MAX_VALUE) ? numBytes : 
> min(until, numBytes){}}}; {{getChar}} returns {{codePointFrom(charIndex)}} 
> directly (the existing {{Objects.checkIndex}} guarantees the index is in 
> range).
>  * Non-ASCII and invalid-UTF-8 strings have the flag set to {{NOT_ASCII}} 
> when warmed, so they take the unchanged scan path.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to