marin-ma opened a new pull request, #13172:
URL: https://github.com/apache/gluten/pull/13172

   ## What changes were proposed in this pull request?
   
   Moves `tools/scripts/gen-function-support-docs.py` from Spark 3.5 to Spark 
4.1 and regenerates
   `docs/velox-backend-{scalar,aggregate,window,generator}-function-support.md` 
against Spark 4.1.1.
   The status describes `spark.sql.ansi.enabled=false`; in ANSI mode Gluten 
falls back to vanilla
   Spark (`spark.gluten.sql.ansiFallback.enabled`).
   
   Script:
   - Replace the embedded FunctionRegistry mapping with Spark 4.1.1's and parse 
its new entry forms
     (multi-line `expressionBuilder`, `<Class>.registryEntry`).
   - Resolve ExpressionBuilder class names to the expression classes they build 
before checking
     Gluten's `ExpressionMappings`. Builder-registered functions (`hour`, 
`minute`, `second`,
     `date_part`, `make_timestamp_*`) were otherwise marked unsupported without 
any test evidence.
   - Run the "not in ExpressionMappings" check for aggregate and generator 
functions too, skipping
     RuntimeReplaceable classes (queried from the JVM) that Spark rewrites 
before Gluten sees them.
   - Parse two fallback reasons the script ignored so far: `Not supported to 
transform StaticInvoke`
     (Spark 4.x implements `encode`, `decode`, `make_timestamp*` with DATE/TIME 
arguments, ... as
     StaticInvoke; a table maps object/method to the function and the run warns 
about new pairs)
     and `Not supported to transform Invoke with function: 
invoke(<X>Evaluator(` (`parse_url`,
     `schema_of_json`, `xpath_*`), mapped back to the function through the 
expression class name.
   - Follow known RuntimeReplaceable rewrites to another documented function
     (`median`/`percentile_cont` -> `percentile`, `zeroifnull` -> `coalesce`, 
`nullifzero` -> `nullif`).
   - Mark cast aliases (`time` for TimeType) unsupported when the log reports 
the type itself as
     unsupported; the aliases are read from the registry literal.
   - Resolve generator class names through Spark's function list instead of a 
hard-coded set, so
     table-valued generators such as `sql_keywords` are reported.
   - Add the Spark 4.1 `variant_funcs` and `st_funcs` scalar groups.
   - Run eight more function suites (URL, JSON, CSV, XML, interval, ST, 
DataFrame window functions,
     table-valued functions).
   - Use the `spark-4.1`/`java-17`/`scala-2.13` profiles and the 
`gluten-ut/spark41` test log.
   - Run the suites with `SPARK_ANSI_SQL_MODE=false`. Spark 4.x enables ANSI 
mode by default and
     Gluten falls back the whole plan in ANSI mode, which hides every 
per-function validation.
   - Run nine SQL query test files that the Spark 4.x settings exclude from CI 
for known result
     gaps but that cover documented functions (`array`, `cast`, `interval`, 
`literals`, `mode`,
     `try_arithmetic`, `try_element_at`, 
`typeCoercion/native/stringCastAndExpressions`, `window`).
   - Deep-copy `KNOWN_RESTRICTIONS`, fix a discarded set union, add 
`regexp_instr` restrictions.
   
   `GlutenSQLQueryTestSuite` (spark40, spark41): two system properties set only 
by the doc generator,
   `gluten.test.sqlQueryTestSuite.ansiEnabled=false` to turn off the ANSI mode 
forced for regular
   test cases (golden files assume ANSI on), and 
`gluten.test.sqlQueryTestSuite.extraTests` to run
   additional files. CI sets neither, so CI behavior is unchanged.
   
   Docs: `docs/developers/UpdateFunctionSupportDocs.md` describes how to move 
the generator to the
   next Spark version and how to audit the generated status; 
`tools/scripts/diff-function-support-docs.py`
   summarizes status changes against a git revision.
   
   ## Function support status changes (Spark 3.5 → Spark 4.1)
   
   | Category  | Spark 3.5 (before)      | Spark 4.1 (after)       |
   |-----------|-------------------------|-------------------------|
   | Scalar    | 357 total, 247 S, 28 PS | 433 total, 261 S, 31 PS |
   | Aggregate | 62 total, 52 S, 1 PS    | 76 total, 52 S, 1 PS    |
   | Window    | 9 total, 9 S            | 9 total, 9 S            |
   | Generator | 7 total, 7 S            | 9 total, 7 S            |
   
   Existing functions whose status changed:
   - Scalar, unsupported → supported: `convert_timezone`, `date_part`, 
`datepart`, `ln`, `radians`,
     `sin`
   - Scalar, unsupported → partially supported: `make_timestamp_ltz`, 
`make_timestamp_ntz`
     (DATE and TIME arguments go through a StaticInvoke Gluten does not 
transform)
   - Scalar, supported → partially supported: `make_timestamp` (same reason),
     `unix_timestamp` (Velox has no `unix_timestamp(DATE, VARCHAR)` signature)
   - Scalar, restriction text only: `regexp_instr` split into "Group index 
ignored" and
     "Lookaround unsupported"; `dayname`, `monthname` dropped the "Spark 4.0+" 
note
   - Aggregate, unsupported → supported: `bitmap_construct_agg`
   - Aggregate, supported → unsupported: `median` (RuntimeReplaceable to 
`percentile`, which Velox
     rejects)
   
   New functions in Spark 4.1:
   - Scalar, supported (5): `<<`, `>>`, `nullifzero`, `randstr`, `zeroifnull`
   - Scalar, unsupported (69): `>>>`, `approx_top_k_estimate`, `collate`, 
`collation`,
     `current_time`, `from_avro`, `from_protobuf`, `from_xml`, `is_valid_utf8`, 
`is_variant_null`,
     `kll_sketch_*` (15), `make_time`, `make_valid_utf8`, `parse_json`, 
`quote`, `schema_of_avro`,
     `schema_of_variant`, `schema_of_variant_agg`, `schema_of_xml`, 
`session_user`, `st_asbinary`,
     `st_geogfromwkb`, `st_geomfromwkb`, `st_setsrid`, `st_srid`, 
`theta_difference`,
     `theta_intersection`, `theta_sketch_estimate`, `theta_union`, `time`, 
`time_diff`,
     `time_trunc`, `to_avro`, `to_protobuf`, `to_time`, `to_variant_object`, 
`to_xml`,
     `try_make_interval`, `try_make_timestamp`, `try_make_timestamp_ltz`, 
`try_make_timestamp_ntz`,
     `try_mod`, `try_parse_json`, `try_parse_url`, `try_reflect`, 
`try_to_date`, `try_to_time`,
     `try_url_decode`, `try_validate_utf8`, `try_variant_get`, `uniform`, 
`validate_utf8`,
     `variant_explode`, `variant_explode_outer`, `variant_get`
   - Aggregate, unsupported (14): `approx_top_k`, `approx_top_k_accumulate`,
     `approx_top_k_combine`, `bitmap_and_agg`, `kll_sketch_agg_bigint`, 
`kll_sketch_agg_double`,
     `kll_sketch_agg_float`, `listagg`, `percentile_cont`, `percentile_disc`, 
`string_agg`,
     `theta_intersection_agg`, `theta_sketch_agg`, `theta_union_agg`
   - Generator, unsupported (2): `collations`, `sql_keywords`
   
   Notes for reviewers:
   - `convert_timezone` (#12979), `bitmap_construct_agg` (#12142) and 
`regexp_instr` (#12697)
     gained native support after the docs were last regenerated. `sin`, `ln`, 
`radians` have been
     mapped for a long time; the committed entries were stale. `date_part`, 
`datepart` were wrongly
     unsupported due to the ExpressionBuilder issue fixed here.
   - Every "supported" row was cross-checked against Gluten's 
`ExpressionMappings` (incl. shims),
     the function names Velox/Gluten register natively, and the test log. That 
audit is what
     surfaced the StaticInvoke, Invoke, RuntimeReplaceable and cast-alias 
handling above: without
     it `encode`, `decode`, `parse_url`, `median`, `percentile_cont`, `time`, 
the `approx_top_k*`
     aggregates and `collations` showed as supported. `parse_url` and `time` 
were already
     unsupported in the Spark 3.5 docs; the new handling keeps them so with log 
evidence.
   - The script marks a function supported when no fallback is logged for it. 
`randstr`, `<<` and
     `>>` have no native test coverage; their status rests on Velox registering 
the function.
   - Test coverage was compared against the Spark 3.5 run: the DataFrame suites 
run at least as
     many tests on 4.1, and with the nine extra SQL files every 3.5 file with 
function coverage
     runs on 4.1 as well.
   
   ## How was this patch tested?
   
   Ran the generator against a Spark 4.1.1 source build with the Spark 4.1 
Gluten build
   (1441 tests passed, 54 golden-file failures expected from the non-ANSI 
setting and the
   deliberately included extra files; the script only reads fallback reasons). 
Verified the
   Spark 3.5 literal extraction reproduces the previous mapping byte for byte, 
that the spark41
   GlutenSQLQueryTestSuite passes Scalastyle and runs the extra files, and 
cross-checked every
   "supported" row against the Velox function registry.
   
   ## Was this patch authored or co-authored using generative AI tooling?
   
   Generated-by: Claude claude-fable-5-1
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to