marin-ma opened a new pull request, #13172:
URL: https://github.com/apache/gluten/pull/13172
## What changes were proposed in this pull request?
Moves `tools/scripts/gen-function-support-docs.py` from Spark 3.5 to Spark
4.1 and regenerates
`docs/velox-backend-{scalar,aggregate,window,generator}-function-support.md`
against Spark 4.1.1.
The status describes `spark.sql.ansi.enabled=false`; in ANSI mode Gluten
falls back to vanilla
Spark (`spark.gluten.sql.ansiFallback.enabled`).
Script:
- Replace the embedded FunctionRegistry mapping with Spark 4.1.1's and parse
its new entry forms
(multi-line `expressionBuilder`, `<Class>.registryEntry`).
- Resolve ExpressionBuilder class names to the expression classes they build
before checking
Gluten's `ExpressionMappings`. Builder-registered functions (`hour`,
`minute`, `second`,
`date_part`, `make_timestamp_*`) were otherwise marked unsupported without
any test evidence.
- Run the "not in ExpressionMappings" check for aggregate and generator
functions too, skipping
RuntimeReplaceable classes (queried from the JVM) that Spark rewrites
before Gluten sees them.
- Parse two fallback reasons the script ignored so far: `Not supported to
transform StaticInvoke`
(Spark 4.x implements `encode`, `decode`, `make_timestamp*` with DATE/TIME
arguments, ... as
StaticInvoke; a table maps object/method to the function and the run warns
about new pairs)
and `Not supported to transform Invoke with function:
invoke(<X>Evaluator(` (`parse_url`,
`schema_of_json`, `xpath_*`), mapped back to the function through the
expression class name.
- Follow known RuntimeReplaceable rewrites to another documented function
(`median`/`percentile_cont` -> `percentile`, `zeroifnull` -> `coalesce`,
`nullifzero` -> `nullif`).
- Mark cast aliases (`time` for TimeType) unsupported when the log reports
the type itself as
unsupported; the aliases are read from the registry literal.
- Resolve generator class names through Spark's function list instead of a
hard-coded set, so
table-valued generators such as `sql_keywords` are reported.
- Add the Spark 4.1 `variant_funcs` and `st_funcs` scalar groups.
- Run eight more function suites (URL, JSON, CSV, XML, interval, ST,
DataFrame window functions,
table-valued functions).
- Use the `spark-4.1`/`java-17`/`scala-2.13` profiles and the
`gluten-ut/spark41` test log.
- Run the suites with `SPARK_ANSI_SQL_MODE=false`. Spark 4.x enables ANSI
mode by default and
Gluten falls back the whole plan in ANSI mode, which hides every
per-function validation.
- Run nine SQL query test files that the Spark 4.x settings exclude from CI
for known result
gaps but that cover documented functions (`array`, `cast`, `interval`,
`literals`, `mode`,
`try_arithmetic`, `try_element_at`,
`typeCoercion/native/stringCastAndExpressions`, `window`).
- Deep-copy `KNOWN_RESTRICTIONS`, fix a discarded set union, add
`regexp_instr` restrictions.
`GlutenSQLQueryTestSuite` (spark40, spark41): two system properties set only
by the doc generator,
`gluten.test.sqlQueryTestSuite.ansiEnabled=false` to turn off the ANSI mode
forced for regular
test cases (golden files assume ANSI on), and
`gluten.test.sqlQueryTestSuite.extraTests` to run
additional files. CI sets neither, so CI behavior is unchanged.
Docs: `docs/developers/UpdateFunctionSupportDocs.md` describes how to move
the generator to the
next Spark version and how to audit the generated status;
`tools/scripts/diff-function-support-docs.py`
summarizes status changes against a git revision.
## Function support status changes (Spark 3.5 → Spark 4.1)
| Category | Spark 3.5 (before) | Spark 4.1 (after) |
|-----------|-------------------------|-------------------------|
| Scalar | 357 total, 247 S, 28 PS | 433 total, 261 S, 31 PS |
| Aggregate | 62 total, 52 S, 1 PS | 76 total, 52 S, 1 PS |
| Window | 9 total, 9 S | 9 total, 9 S |
| Generator | 7 total, 7 S | 9 total, 7 S |
Existing functions whose status changed:
- Scalar, unsupported → supported: `convert_timezone`, `date_part`,
`datepart`, `ln`, `radians`,
`sin`
- Scalar, unsupported → partially supported: `make_timestamp_ltz`,
`make_timestamp_ntz`
(DATE and TIME arguments go through a StaticInvoke Gluten does not
transform)
- Scalar, supported → partially supported: `make_timestamp` (same reason),
`unix_timestamp` (Velox has no `unix_timestamp(DATE, VARCHAR)` signature)
- Scalar, restriction text only: `regexp_instr` split into "Group index
ignored" and
"Lookaround unsupported"; `dayname`, `monthname` dropped the "Spark 4.0+"
note
- Aggregate, unsupported → supported: `bitmap_construct_agg`
- Aggregate, supported → unsupported: `median` (RuntimeReplaceable to
`percentile`, which Velox
rejects)
New functions in Spark 4.1:
- Scalar, supported (5): `<<`, `>>`, `nullifzero`, `randstr`, `zeroifnull`
- Scalar, unsupported (69): `>>>`, `approx_top_k_estimate`, `collate`,
`collation`,
`current_time`, `from_avro`, `from_protobuf`, `from_xml`, `is_valid_utf8`,
`is_variant_null`,
`kll_sketch_*` (15), `make_time`, `make_valid_utf8`, `parse_json`,
`quote`, `schema_of_avro`,
`schema_of_variant`, `schema_of_variant_agg`, `schema_of_xml`,
`session_user`, `st_asbinary`,
`st_geogfromwkb`, `st_geomfromwkb`, `st_setsrid`, `st_srid`,
`theta_difference`,
`theta_intersection`, `theta_sketch_estimate`, `theta_union`, `time`,
`time_diff`,
`time_trunc`, `to_avro`, `to_protobuf`, `to_time`, `to_variant_object`,
`to_xml`,
`try_make_interval`, `try_make_timestamp`, `try_make_timestamp_ltz`,
`try_make_timestamp_ntz`,
`try_mod`, `try_parse_json`, `try_parse_url`, `try_reflect`,
`try_to_date`, `try_to_time`,
`try_url_decode`, `try_validate_utf8`, `try_variant_get`, `uniform`,
`validate_utf8`,
`variant_explode`, `variant_explode_outer`, `variant_get`
- Aggregate, unsupported (14): `approx_top_k`, `approx_top_k_accumulate`,
`approx_top_k_combine`, `bitmap_and_agg`, `kll_sketch_agg_bigint`,
`kll_sketch_agg_double`,
`kll_sketch_agg_float`, `listagg`, `percentile_cont`, `percentile_disc`,
`string_agg`,
`theta_intersection_agg`, `theta_sketch_agg`, `theta_union_agg`
- Generator, unsupported (2): `collations`, `sql_keywords`
Notes for reviewers:
- `convert_timezone` (#12979), `bitmap_construct_agg` (#12142) and
`regexp_instr` (#12697)
gained native support after the docs were last regenerated. `sin`, `ln`,
`radians` have been
mapped for a long time; the committed entries were stale. `date_part`,
`datepart` were wrongly
unsupported due to the ExpressionBuilder issue fixed here.
- Every "supported" row was cross-checked against Gluten's
`ExpressionMappings` (incl. shims),
the function names Velox/Gluten register natively, and the test log. That
audit is what
surfaced the StaticInvoke, Invoke, RuntimeReplaceable and cast-alias
handling above: without
it `encode`, `decode`, `parse_url`, `median`, `percentile_cont`, `time`,
the `approx_top_k*`
aggregates and `collations` showed as supported. `parse_url` and `time`
were already
unsupported in the Spark 3.5 docs; the new handling keeps them so with log
evidence.
- The script marks a function supported when no fallback is logged for it.
`randstr`, `<<` and
`>>` have no native test coverage; their status rests on Velox registering
the function.
- Test coverage was compared against the Spark 3.5 run: the DataFrame suites
run at least as
many tests on 4.1, and with the nine extra SQL files every 3.5 file with
function coverage
runs on 4.1 as well.
## How was this patch tested?
Ran the generator against a Spark 4.1.1 source build with the Spark 4.1
Gluten build
(1441 tests passed, 54 golden-file failures expected from the non-ANSI
setting and the
deliberately included extra files; the script only reads fallback reasons).
Verified the
Spark 3.5 literal extraction reproduces the previous mapping byte for byte,
that the spark41
GlutenSQLQueryTestSuite passes Scalastyle and runs the extra files, and
cross-checked every
"supported" row against the Velox function registry.
## Was this patch authored or co-authored using generative AI tooling?
Generated-by: Claude claude-fable-5-1
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]