[
https://issues.apache.org/jira/browse/SPARK-59513?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59513:
-----------------------------------
Labels: pull-request-available (was: )
> assertDataFrameEqual silently ignores extra fields in Row values
> ----------------------------------------------------------------
>
> Key: SPARK-59513
> URL: https://issues.apache.org/jira/browse/SPARK-59513
> Project: Spark
> Issue Type: Bug
> Components: PySpark
> Affects Versions: 5.0.0
> Reporter: Akshay Thorat
> Priority: Major
> Labels: pull-request-available
>
> h2. Problem
> The Row comparator uses zip without a length check, so trailing fields are
> ignored. Supported list-of-Row inputs do not undergo a DataFrame schema
> comparison, allowing unequal results to pass assertions.
> h2. Reproduction
> Reproduced on upstream master 39776477a3d (PySpark 5.0.0.dev0), Python
> 3.10.11. No SparkSession is required.
> {code:python}
> from pyspark.sql import Row
> from pyspark.testing import assertDataFrameEqual
> assertDataFrameEqual([Row(x=1)], [Row(x=1, y=2)])
> # Returns without raising
> {code}
> h2. Expected behavior
> Raise PySparkAssertionError with condition DIFFERENT_ROWS.
> h2. Proposed fix and verification
> Check Row lengths before recursively comparing field values. Tests cover both
> argument directions, empty Rows, nested Rows, arrays, maps, and both
> row-order settings. Regression tests fail before the fix and pass afterward.
> Python compilation, custom-error checks and Ruff checks pass.
> Pull request: https://github.com/apache/spark/pull/58789
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]