Yaoxuan Wu created SPARK-60078:
----------------------------------

             Summary: Column.outer() binds to a column of the inner plan after 
a rename, giving wrong results
                 Key: SPARK-60078
                 URL: https://issues.apache.org/jira/browse/SPARK-60078
             Project: Spark
          Issue Type: Bug
          Components: SQL
    Affects Versions: 4.1.1, 4.2.0, 4.0.0
            Reporter: Yaoxuan Wu


If the subquery DataFrame renamed a column ({{{}toDF{}}}, 
{{{}withColumnRenamed{}}}, {{{}select(col.alias(...)){}}}) and the outer 
DataFrame has a column with the old name, an unqualified 
{{F.col(old_name).outer()}} is bound to the subquery's own pre-rename column. 
The correlation predicate becomes a predicate on the subquery alone and the 
query returns wrong results without any error.

 
{code:java}
from pyspark.sql import SparkSession
spark = SparkSession.builder.getOrCreate()
from pyspark.sql import functions as F
a = spark.createDataFrame([(1,), (2,), (3,)], ['k'])
b = spark.createDataFrame([(1,), (5,)], ['k']).toDF('k_r')
pred = F.col('k_r') == F.col('k').outer()      # intended: b.k_r = 
a.kprint(a.lateralJoin(b.where(pred)).count())                                 
# 6, expected 1
print(a.where(b.where(pred).exists()).count())                              # 
3, expected 1
print(a.select(b.where(pred).select(F.count('*')).scalar()).collect())      # 
2, 2, 2; expected 1, 0, 0 {code}
Analyzed plan (4.2.0):
{code:java}
LateralJoin lateral-subquery#102 [], Inner
:  +- Project [k_r#101L]
:     +- Filter (k_r#101L = k#1L)
:        +- Project [k#1L AS k_r#101L, k#1L]
:           +- LogicalRDD [k#1L], false
+- LogicalRDD [k#0L], false {code}
k is resolved as k#1 of the subquery, not as the outer k#0.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to