[ 
https://issues.apache.org/jira/browse/SPARK-59506?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SPARK-59506:
-----------------------------------
    Labels: pull-request-available  (was: )

> Avoid per-call regex Matcher allocation in SparkDateTimeUtils.getZoneId
> -----------------------------------------------------------------------
>
>                 Key: SPARK-59506
>                 URL: https://issues.apache.org/jira/browse/SPARK-59506
>             Project: Spark
>          Issue Type: Improvement
>          Components: SQL
>    Affects Versions: 4.1.0
>            Reporter: David Mollitor
>            Priority: Minor
>              Labels: pull-request-available
>
> {{SparkDateTimeUtils.getZoneId(String)}} normalizes the pre-Spark-3.0 
> single-digit UTC-offset forms ({{{}({+}|{-})h:mm{-}{+}{}}} and 
> {{{}(|)hh:m{}}}) by running two regexes on every call:
> {code:scala}
> final val singleHourTz = Pattern.compile("(\+|\-)(\d):")
> final val singleMinuteTz = Pattern.compile("(\+|\-)(\d\d):(\d)$")
> def getZoneId(timeZoneId: String): ZoneId = {
>   var formattedZoneId = 
> singleHourTz.matcher(timeZoneId).replaceFirst("$10$2:")
>   formattedZoneId = 
> singleMinuteTz.matcher(formattedZoneId).replaceFirst("$1$2:0$3")
>   ZoneId.of(formattedZoneId, ZoneId.SHORT_IDS)
> }
> {code}
> {{Pattern.matcher()}} allocates a fresh {{Matcher}} on every call. 
> {{getZoneId}} sits on the per-row timestamp-parse hot path 
> ({{{}UnivocityParser{}}} -> {{TimestampFormatter.parse}} -> 
> {{stringToTimestamp}} -> {{parseTimestampString}} -> {{{}getZoneId{}}}), so a 
> timestamp column allocates two throwaway Matchers *per row* to run a 
> normalization that changes nothing for the common inputs ({{{}Z{}}}, 
> {{{}+07:30{}}}, named zones, {{{}UTC{}}}).
> Also, the current implementation runs two regexes over the string – each
> {{Matcher.replaceFirst}} scans the input with {{{}find(){}}}, so a call 
> walked the
> string twice (and allocated a {{Matcher}} per scan). This proposed 
> replacement makes a single linear pass for the hour rule; the minute rule is 
> end-anchored, so it is just a constant-time check of the last few characters 
> rather than a second scan.
> h3. Profiling evidence
> JFR profiling of a CSV timestamp-parsing workload found {{getZoneId -> 
> Pattern.matcher}} to be the single largest {{java.util.regex.Matcher}} 
> allocation site: ~97% of all Matcher allocation and roughly 8% of total 
> sampled allocation pressure in the run.
> h3. Proposed change
> Replace the two regexes with an allocation-free string walk that applies the 
> same
> normalization only when a legacy single-digit offset is actually present, and 
> returns the
> input string unchanged (same reference) otherwise:
>  * {{({+}|{-})h:mm{-}{+}}} -> {{(|)0h:mm}} (pad the hour; first occurrence 
> only)
>  * {{({+}|{-})hh:m{-}{+}}} -> {{(|)hh:0m}} (pad the minute; only when it ends 
> the string)
> The {{Pattern}} fields are removed. Behavior is preserved exactly, including 
> the legacy-format normalization and the {{INVALID_TIMEZONE}} error raised for 
> invalid zones. The change also benefits the many one-time callers 
> ({{{}CSVOptions{}}}/{{{}JSONOptions{}}}/{{{}XmlOptions{}}}, session-timezone 
> resolution, etc.).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to