Hi all, We've had some great input and discussion in the DECFLOAT SPIP document (https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E/edit?tab=t.0#heading=h.zarpams0nxzq), so I'm writing an update here to keep the DISCUSS thread active and invite further feedback and discussion regarding this SPIP.
It's really important to separate the irreversible decisions from the reversible ones. The one-way door here is the value domain - DECFLOAT should carry IEEE 754's non-finite values (+Inf, -Inf, NaN, signed zero), the same way DOUBLE already does. Digit count (e.g. 34 vs 38) and arithmetic implementation can easily be changed later - for example, digits can always be widened later and retain full backwards compatibility; also, the execution kernel can change under a fixed IEEE contract. However, specials are at the opposite end - we can't retrofit them onto a shipped finite-only type without breaking it. Why this matters for big data: decimal pipelines produce unbounded and undefined quantities as legitimate business states, not data bugs (e.g. an uncapped credit line, a threshold that means "no maximum," a magnitude that has run past any finite bound). IEEE already names & defines these precisely, and DOUBLE carries them today. DECFLOAT is the base-10 DOUBLE, its domain should match accordingly. A finite-only decimal type can't hold those values, so applications disguise them (almost always as NULL). However, NULL in Spark means missing/unknown, so "unbounded" and "never loaded" would become the same value, and SUM, COUNT, range filters, etc. could no longer telling them apart. The loss here runs one way, we can always collapse Inf/NaN to NULL if we want SQL null semantics, but we can never rebuild meaning already flattened into NULL. This reaches past Spark too, Parquet's decimal-float work and existing C++/Rust/Python engines already emit Inf/NaN, so a finite-only type can't round-trip that data faithfully across different languages / engines. Let's agree that DECFLOAT's domain is IEEE 754's, non-finite values included. Without this, DECFLOAT isn't a better DECIMAL - it's a worse DOUBLE. Best, Uroš On 2026/08/18 16:00:13 Uroš Bojanić wrote: > Hi all, > > I would like to start a discussion on the SPIP to add a new Spark SQL data > type: DECFLOAT (IEEE 754 decimal64 / decimal128), for base-10 floating-point > decimals with per-value exponents. > > JIRA ID: https://issues.apache.org/jira/browse/SPARK-58820 > > Brief summary: DECFLOAT is a new IEEE 754 decimal floating-point type > (decimal64 / decimal128, i.e. DECFLOAT(16) / DECFLOAT(34)) that closes the > gap between DECIMAL, which caps at precision 38 with a fixed per-column > scale, and DOUBLE, whose binary rounding makes fractions like 0.1 inexact: > each value keeps its own exponent, arithmetic is decimal, and it supports > signed zero, Inf, and NaN. It is additive, round-trips through a proposed > Parquet logical type for cross-engine interop, and would ship config-gated > during incubation like TIME, leaving existing DECIMAL / DOUBLE behavior > unchanged. > > Additional information is available in the SPIP document: > https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E > > Please provide your feedback on the approach, scope, and the proposal > described in the document. > > Thank you! > > Best, > Uroš > > --------------------------------------------------------------------- > To unsubscribe e-mail: [email protected] > > --------------------------------------------------------------------- To unsubscribe e-mail: [email protected]
