[
https://issues.apache.org/jira/browse/SPARK-57925?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18097299#comment-18097299
]
Marcus Lin commented on SPARK-57925:
------------------------------------
If this is still available, I would like to work on it!
> Add configurable default scale for Oracle bare NUMBER columns
> -------------------------------------------------------------
>
> Key: SPARK-57925
> URL: https://issues.apache.org/jira/browse/SPARK-57925
> Project: Spark
> Issue Type: Improvement
> Components: SQL
> Affects Versions: 4.2.0
> Reporter: jiangxintong
> Priority: Major
>
> h2. Problem
> Oracle JDBC driver returns precision=0 and scale=-127 for bare NUMBER columns
> (no precision/scale declared). Spark's OracleDialect hardcodes this to
> DecimalType(38, 10), causing silent precision loss for values with more than
> 10 decimal places.
> {code:java}
> // OracleDialect.scala:147-162
> case 0 => Option(DecimalType(DecimalType.MAX_PRECISION, 10)) // hardcoded
> scale=10
> case _ if scale == -127L => Option(DecimalType(DecimalType.MAX_PRECISION, 10))
> {code}
> Example: A NUMBER column storing 3.141592653589793 is truncated to
> 3.1415926536 when read via Spark JDBC.
> h2. Related Issues
> * SPARK-23370: Resolved/Incomplete — closed as "the problem is not
> completely described," not Won't Fix. No concrete fix was proposed.
> * #56738: GitHub issue reporting the same problem.
> * apache/spark PR #8780 (2015): The original hardcoding.
> h2. Proposed Solution
> Add a new SQL configuration: spark.sql.oracle.jdbc.number.defaultScale
> (default: 10, backward compatible). Config name TBD during review.
> In OracleDialect.getCatalystType, read this configuration instead of
> hardcoding 10. This is an opt-in escape hatch — the default behavior is
> unchanged, but users who need higher precision can set:
> {code:sql}
> SET spark.sql.oracle.jdbc.number.defaultScale = 15;
> {code}
> h2. Why Configuration Instead of Changing the Default
> The Oracle JDBC driver returns scale=-127 for bare NUMBER, which means
> "floating point with unknown scale." There is no "real" scale to read. The
> original PR (apache/spark PR #8780, 2015) hardcoded 10 as a pragmatic
> default. A configuration is a clean escape hatch that preserves backward
> compatibility while allowing users to opt into higher precision.
> h2. Why SQLConf Instead of Per-Read JDBC Option
> getCatalystType(sqlType, typeName, size, md) signature does not have access
> to per-read JDBC options. This is the same pattern used by
> MsSqlServerDialect, DB2Dialect, MySQLDialect, and PostgresDialect, all of
> which read SQLConf.get for dialect-specific behavior.
> h2. Validation
> The configuration value must be in [0, 38] (MAX_PRECISION). Values outside
> this range should throw IllegalArgumentException.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]