[ 
https://issues.apache.org/jira/browse/SPARK-60010?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SPARK-60010:
-----------------------------------
    Labels: pull-request-available  (was: )

> Reading a Parquet DECIMAL column with an integer schema returns the raw 
> unscaled integer instead of honoring the decimal scale
> ------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: SPARK-60010
>                 URL: https://issues.apache.org/jira/browse/SPARK-60010
>             Project: Spark
>          Issue Type: Bug
>          Components: SQL
>    Affects Versions: 4.1.1
>            Reporter: Brian Wu
>            Priority: Major
>              Labels: pull-request-available
>
> When a Parquet column physically stored as DECIMAL is read with an integer 
> read-schema, Spark's behavior is unexpected: for low-precision decimals it 
> silently returns the raw unscaled stored integer (ignoring the scale), while 
> for higher-precision decimals it throws a type-mismatch error.
>  
> Reproduction:
> {code:java}
> import org.apache.spark.sql.types._
> val intSchema = StructType(Seq(StructField("v", IntegerType)))
> Case A – DECIMAL(9,2), physically stored as INT32:
> spark.sql("SELECT CAST(123.45 AS DECIMAL(9,2)) AS 
> v").write.mode("overwrite").parquet("/tmp/dec92")
> spark.read.schema(intSchema).parquet("/tmp/dec92").show()
> Result - silently returns the unscaled stored value (scale ignored)
> +-----+
> |    v|
> +-----+
> |12345|
> +-----+ 
> Case B – DECIMAL(10,2), physically stored as INT64:
> spark.sql("SELECT CAST(123.45 AS DECIMAL(10,2)) AS 
> v").write.mode("overwrite").parquet("/tmp/dec102")
> spark.read.schema(intSchema).parquet("/tmp/dec102").show()
> Result - throws
> org.apache.spark.SparkException: 
> [FAILED_READ_FILE.PARQUET_COLUMN_DATA_TYPE_MISMATCH] Encountered error while 
> reading file 
> file:///tmp/dec102/part-00000-a59fef00-bb30-4788-a195-2f4c6173117e-c000.snappy.parquet.
>  Data type mismatches when reading Parquet column [v]. Expected Spark type 
> int, actual Parquet type INT64{code}
>  
> [https://github.com/apache/parquet-format/blob/master/LogicalTypes.md#numeric-types]
> Parquet standard does explicitly not define behavior here. However, it seems 
> Spark behavior does not make sense in these cases. Suggest Spark should honor 
> the schema type and not interpret the physical value as unscaled.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to