Thanks for updating the proposal and working through the feedback. I support adding decimal floating-point support to Spark.
The revised proposal now specifies an in-tree, pure-Java implementation with no native runtime dependency. This appears to address the architectural concern behind Dongjoon’s original -1. Dongjoon, could you confirm whether this addresses your concern about having a safe Java execution path? Thanks, On Mon, Sep 21, 2026 at 3:40 PM Hyukjin Kwon <[email protected]> wrote: > Let's separate the implementation specific part from SPIP. For > implementation things, I think we can discuss details in the PR and make > sure, during the reviews, to involve people who had concerns about it. > > On Fri, 18 Sept 2026 at 02:16, Gengliang Wang <[email protected]> wrote: > >> Thanks for the update. I agree with separating the value domain from >> exception handling and the arithmetic implementation. >> >> I support defining DECFLOAT with signed zero, positive and negative >> infinity, and quiet NaN. These are numeric values rather than SQL NULL, and >> ppreserving them enables faithful interchange and round-tripping across >> language and engine boundaries, including between Spark and Python clients. >> A finite-only definition would lose that information and would be difficult >> to extend compatibly later. >> >> Detailed coercion rules, the execution implementation, external APIs, and >> the Parquet encoding still require further review, but they should not >> block agreement on the logical value domain. >> >> With this separation clarified, I support moving the SPIP forward. >> >> Best, >> Gengliang >> >> On Thu, Sep 17, 2026 at 7:44 AM Uroš Bojanić <[email protected]> wrote: >> >>> Hi all, >>> >>> We've had some great input and discussion in the DECFLOAT SPIP document ( >>> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E/edit?tab=t.0#heading=h.zarpams0nxzq), >>> so I'm writing an update here to keep the DISCUSS thread active and invite >>> further feedback and discussion regarding this SPIP. >>> >>> It's really important to separate the irreversible decisions from the >>> reversible ones. The one-way door here is the value domain - DECFLOAT >>> should carry IEEE 754's non-finite values (+Inf, -Inf, NaN, >>> signed zero), the same way DOUBLE already does. Digit count (e.g. 34 vs >>> 38) and arithmetic implementation can easily be changed later - for >>> example, digits can always be widened later and retain full backwards >>> compatibility; also, the execution kernel can change under a fixed IEEE >>> contract. However, specials are at the opposite end - we can't retrofit >>> them onto a shipped finite-only type without breaking it. >>> >>> Why this matters for big data: decimal pipelines produce unbounded and >>> undefined quantities as legitimate business states, not data bugs (e.g. an >>> uncapped credit line, a threshold that means "no maximum," a magnitude that >>> has run past any finite bound). IEEE already names & defines these >>> precisely, and DOUBLE carries them today. DECFLOAT is the base-10 DOUBLE, >>> its domain should match accordingly. >>> >>> A finite-only decimal type can't hold those values, so applications >>> disguise them (almost always as NULL). However, NULL in Spark means >>> missing/unknown, so "unbounded" and "never loaded" would become the same >>> value, and SUM, COUNT, range filters, etc. could no longer telling them >>> apart. The loss here runs one way, we can always collapse Inf/NaN to NULL >>> if we want SQL null semantics, but we can never rebuild meaning already >>> flattened into NULL. This reaches past Spark too, Parquet's decimal-float >>> work and existing C++/Rust/Python engines already emit Inf/NaN, so a >>> finite-only type can't round-trip that data faithfully across different >>> languages / engines. >>> >>> Let's agree that DECFLOAT's domain is IEEE 754's, non-finite values >>> included. Without this, DECFLOAT isn't a better DECIMAL - it's a worse >>> DOUBLE. >>> >>> Best, >>> Uroš >>> >>> On 2026/08/18 16:00:13 Uroš Bojanić wrote: >>> > Hi all, >>> > >>> > I would like to start a discussion on the SPIP to add a new Spark SQL >>> data type: DECFLOAT (IEEE 754 decimal64 / decimal128), for base-10 >>> floating-point decimals with per-value exponents. >>> > >>> > JIRA ID: https://issues.apache.org/jira/browse/SPARK-58820 >>> > >>> > Brief summary: DECFLOAT is a new IEEE 754 decimal floating-point type >>> (decimal64 / decimal128, i.e. DECFLOAT(16) / DECFLOAT(34)) that closes the >>> gap between DECIMAL, which caps at precision 38 with a fixed per-column >>> scale, and DOUBLE, whose binary rounding makes fractions like 0.1 inexact: >>> each value keeps its own exponent, arithmetic is decimal, and it supports >>> signed zero, Inf, and NaN. It is additive, round-trips through a proposed >>> Parquet logical type for cross-engine interop, and would ship config-gated >>> during incubation like TIME, leaving existing DECIMAL / DOUBLE behavior >>> unchanged. >>> > >>> > Additional information is available in the SPIP document: >>> > >>> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E >>> > >>> > Please provide your feedback on the approach, scope, and the proposal >>> described in the document. >>> > >>> > Thank you! >>> > >>> > Best, >>> > Uroš >>> > >>> > --------------------------------------------------------------------- >>> > To unsubscribe e-mail: [email protected] >>> > >>> > >>> >>> --------------------------------------------------------------------- >>> To unsubscribe e-mail: [email protected] >>> >>>
