Hyukjin Kwon created SPARK-58998:
------------------------------------

             Summary: [spark-connect-rust] Drop-in pyspark API fidelity + 
high-level API test coverage
                 Key: SPARK-58998
                 URL: https://issues.apache.org/jira/browse/SPARK-58998
             Project: Spark
          Issue Type: Improvement
          Components: Project Infra
    Affects Versions: 5.0.0
            Reporter: Hyukjin Kwon


Follow-up to SPARK-58987 (which added the ours-only parity gate + coverage 
tooling).

Two findings from the coverage work:
h2. 1. The parity gate covers the wire, not our high-level API
The transport-injection harness swaps only the gRPC stub, so the official suite 
runs the REFERENCE pyspark DataFrame/Session/Column classes and only routes 
bytes through our Rust stub. Our high-level API (spark-connect 
DataFrame/Session/Catalog/... and the pyspark-rs bindings) is therefore NOT 
exercised by it - merged coverage is ~49%, with those modules near 0-20% while 
the wire layer (client/error/retries/channel/bytes_codec) is 90%+.

h2. 2. The drop-in pyspark API diverges from reference pyspark
The skin re-exports the Rust PyO3 classes directly, and their API is 
name-divergent and incomplete vs reference pyspark, e.g.:
* SparkSession: no `conf`, no `read` (only `readStream`); `catalog` is a 
method, not a property
* DataFrame: `where_` not `where`; missing `withColumns`, `hint`, 
`unpivot`/`melt`, `intersectAll`; `head()` takes no count; `createDataFrame` 
accepts only a list of column names (no StructType/DDL)
* Column: `is_null`/`is_not_null`/`get_field` (snake_case) not 
`isNull`/`isNotNull`/`getField`
* GroupedData: no `pivot`
The origin project reached full parity via a skin adapter layer; the fork 
appears to have lost it. `df.where(...)`, `col.isNull()`, `spark.read`, etc. 
would break for users.

h2. Work
* Restore a faithful drop-in skin (adapter wrappers: where_->where, 
is_null->isNull, add 
conf/read/pivot/withColumns/hint/unpivot/intersectAll/head(n)/StructType-createDataFrame,
 catalog as a property).
* Add high-level API tests (Rust golden+e2e for the spark-connect API; a 
drop-in Python exercise, scripts/e2e_wrapper.py, for the pyspark-rs bindings).
* Raise the rust-coverage CI gate (scripts/rust_coverage.sh, added in 
SPARK-58987) to 90%.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to