[
https://issues.apache.org/jira/browse/SPARK-58232?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Ganesha S updated SPARK-58232:
------------------------------
Description:
*Summary*
In DefaultCachedBatchSerializer
(sql/core/.../execution/columnar/InMemoryRelation.scala), the mapping from
selected
output attributes to their ordinals in the cached schema is computed with an
O(n*m)
algorithm: for each selected attribute it rebuilds the full list of
cached-schema
ExprIds and linear-scans it with indexOf. This replaces that with a single
ExprId -> ordinal map and O(1) lookups (overall O(n+m)).
This is a code-quality cleanup, not a performance fix. The computation runs
once per
scan execution on the driver, so the runtime impact is negligible in practice;
the value
is removing a quadratic algorithm and a per-attribute list re-allocation in
favor of the
straightforward map lookup.
*Affected code*
sql/core/src/main/scala/org/apache/spark/sql/execution/columnar/InMemoryRelation.scala
Both conversion methods share the pattern:
* convertCachedBatchToColumnarBatch (vectorized path)
* convertCachedBatchToInternalRow (row path)
{code:scala}
val columnIndices =
selectedAttributes.map(a => cacheAttributes.map(o =>
o.exprId).indexOf(a.exprId)).toArray
{code}
For each of m selected attributes this rebuilds a fresh Seq[ExprId] from all n
cached
attributes
was:
*Summary*
When reading from an in-memory (cached) relation,
`DefaultCachedBatchSerializer` maps
each selected output attribute to its ordinal in the cached schema. It
currently does
this with an `O(n*m)` algorithm that, for every selected attribute, rebuilds
the full
list of cached-schema ExprIds and linear-scans it with indexOf. This is
wasteful time
and transient allocation on wide cached scans and should be a single map lookup.
*Affected code*
sql/core/src/main/scala/org/apache/spark/sql/execution/columnar/InMemoryRelation.scala
Both conversion methods share the pattern:
* convertCachedBatchToColumnarBatch (vectorized path)
* convertCachedBatchToInternalRow (row path)
{code:scala}
val columnIndices =
selectedAttributes.map(a => cacheAttributes.map(o =>
o.exprId).indexOf(a.exprId)).toArray
{code}
For each of m selected attributes this rebuilds a fresh Seq[ExprId] from all n
cached
attributes
> [SQL] Simplify cached-batch column-index resolution with a map lookup
> ----------------------------------------------------------------------
>
> Key: SPARK-58232
> URL: https://issues.apache.org/jira/browse/SPARK-58232
> Project: Spark
> Issue Type: Task
> Components: SQL
> Affects Versions: 4.2.0
> Reporter: Ganesha S
> Priority: Major
>
> *Summary*
> In DefaultCachedBatchSerializer
> (sql/core/.../execution/columnar/InMemoryRelation.scala), the mapping from
> selected
> output attributes to their ordinals in the cached schema is computed with an
> O(n*m)
> algorithm: for each selected attribute it rebuilds the full list of
> cached-schema
> ExprIds and linear-scans it with indexOf. This replaces that with a single
> ExprId -> ordinal map and O(1) lookups (overall O(n+m)).
> This is a code-quality cleanup, not a performance fix. The computation runs
> once per
> scan execution on the driver, so the runtime impact is negligible in
> practice; the value
> is removing a quadratic algorithm and a per-attribute list re-allocation in
> favor of the
> straightforward map lookup.
>
> *Affected code*
> sql/core/src/main/scala/org/apache/spark/sql/execution/columnar/InMemoryRelation.scala
> Both conversion methods share the pattern:
> * convertCachedBatchToColumnarBatch (vectorized path)
> * convertCachedBatchToInternalRow (row path)
> {code:scala}
> val columnIndices =
> selectedAttributes.map(a => cacheAttributes.map(o =>
> o.exprId).indexOf(a.exprId)).toArray
> {code}
> For each of m selected attributes this rebuilds a fresh Seq[ExprId] from all
> n cached
> attributes
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]