Let's separate the implementation specific part from SPIP. For implementation things, I think we can discuss details in the PR and make sure, during the reviews, to involve people who had concerns about it.
On Fri, 18 Sept 2026 at 02:16, Gengliang Wang <[email protected]> wrote: > Thanks for the update. I agree with separating the value domain from > exception handling and the arithmetic implementation. > > I support defining DECFLOAT with signed zero, positive and negative > infinity, and quiet NaN. These are numeric values rather than SQL NULL, and > ppreserving them enables faithful interchange and round-tripping across > language and engine boundaries, including between Spark and Python clients. > A finite-only definition would lose that information and would be difficult > to extend compatibly later. > > Detailed coercion rules, the execution implementation, external APIs, and > the Parquet encoding still require further review, but they should not > block agreement on the logical value domain. > > With this separation clarified, I support moving the SPIP forward. > > Best, > Gengliang > > On Thu, Sep 17, 2026 at 7:44 AM Uroš Bojanić <[email protected]> wrote: > >> Hi all, >> >> We've had some great input and discussion in the DECFLOAT SPIP document ( >> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E/edit?tab=t.0#heading=h.zarpams0nxzq), >> so I'm writing an update here to keep the DISCUSS thread active and invite >> further feedback and discussion regarding this SPIP. >> >> It's really important to separate the irreversible decisions from the >> reversible ones. The one-way door here is the value domain - DECFLOAT >> should carry IEEE 754's non-finite values (+Inf, -Inf, NaN, >> signed zero), the same way DOUBLE already does. Digit count (e.g. 34 vs >> 38) and arithmetic implementation can easily be changed later - for >> example, digits can always be widened later and retain full backwards >> compatibility; also, the execution kernel can change under a fixed IEEE >> contract. However, specials are at the opposite end - we can't retrofit >> them onto a shipped finite-only type without breaking it. >> >> Why this matters for big data: decimal pipelines produce unbounded and >> undefined quantities as legitimate business states, not data bugs (e.g. an >> uncapped credit line, a threshold that means "no maximum," a magnitude that >> has run past any finite bound). IEEE already names & defines these >> precisely, and DOUBLE carries them today. DECFLOAT is the base-10 DOUBLE, >> its domain should match accordingly. >> >> A finite-only decimal type can't hold those values, so applications >> disguise them (almost always as NULL). However, NULL in Spark means >> missing/unknown, so "unbounded" and "never loaded" would become the same >> value, and SUM, COUNT, range filters, etc. could no longer telling them >> apart. The loss here runs one way, we can always collapse Inf/NaN to NULL >> if we want SQL null semantics, but we can never rebuild meaning already >> flattened into NULL. This reaches past Spark too, Parquet's decimal-float >> work and existing C++/Rust/Python engines already emit Inf/NaN, so a >> finite-only type can't round-trip that data faithfully across different >> languages / engines. >> >> Let's agree that DECFLOAT's domain is IEEE 754's, non-finite values >> included. Without this, DECFLOAT isn't a better DECIMAL - it's a worse >> DOUBLE. >> >> Best, >> Uroš >> >> On 2026/08/18 16:00:13 Uroš Bojanić wrote: >> > Hi all, >> > >> > I would like to start a discussion on the SPIP to add a new Spark SQL >> data type: DECFLOAT (IEEE 754 decimal64 / decimal128), for base-10 >> floating-point decimals with per-value exponents. >> > >> > JIRA ID: https://issues.apache.org/jira/browse/SPARK-58820 >> > >> > Brief summary: DECFLOAT is a new IEEE 754 decimal floating-point type >> (decimal64 / decimal128, i.e. DECFLOAT(16) / DECFLOAT(34)) that closes the >> gap between DECIMAL, which caps at precision 38 with a fixed per-column >> scale, and DOUBLE, whose binary rounding makes fractions like 0.1 inexact: >> each value keeps its own exponent, arithmetic is decimal, and it supports >> signed zero, Inf, and NaN. It is additive, round-trips through a proposed >> Parquet logical type for cross-engine interop, and would ship config-gated >> during incubation like TIME, leaving existing DECIMAL / DOUBLE behavior >> unchanged. >> > >> > Additional information is available in the SPIP document: >> > >> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E >> > >> > Please provide your feedback on the approach, scope, and the proposal >> described in the document. >> > >> > Thank you! >> > >> > Best, >> > Uroš >> > >> > --------------------------------------------------------------------- >> > To unsubscribe e-mail: [email protected] >> > >> > >> >> --------------------------------------------------------------------- >> To unsubscribe e-mail: [email protected] >> >>
