Dongjoon,

You make many valid points, and we’ll address them.
Some I do not understand however, and as the author of the Decimal128 PR let me 
ask for clarification:
It is my understanding that it is frowned upon to roll in “unused APIs” (i.e. 
dead code).
The PR remains in draft mode because this SPIP has not passed yet. The PR’s 
initial purpose is to support THIS SPIP.
By the same token, there ought not be a DecFloatType without this SPIP.
Are you advocating:
a) Roll in this library as a precondition to approve the SPIP
b) Make it clear that this is the specific library being used,
c) Sssert that there will be “some” java native library. Whether that is 
BigDecimal based, libbid, or something else - possibly home grown.

I’m also confused by the asserted dependency on Parquet and the transport 
layer. Aside from the obvious chicken-egg problem I wonder why the in-memory 
layout used by the DECFLOAT type in Spark must be aligned with the on-disk 
layout of Parquet.
TIMESTAMP nano being one example where that is not case. Certainly it is “nice” 
to just copy through, but, in my opinion semantics are way more important. 
Which is why we have been focusing on this.

Cheers
Serge

On Sep 22, 2026, at 8:29 AM, Thomas Kissinger via dev <[email protected]> 
wrote:

Hi all,

Thanks, Uroš, for revising the proposal and for all the work invested. I 
support adding a decimal-floating-point type to Spark, and I agree with 
Dongjoon that the proposal is not yet ready for another vote.

My concern is that IEEE BID seems to be a predetermined choice rather than the 
result of comparing alternatives against Spark's requirements. IEEE BID is 
standardized, but its implementation ecosystem is very limited: neither Java 
nor Python provides a native BID value type. The proposed libbid-based 
arithmetic stops at 34 digits, leaving the established 38-digit SQL precision 
unsupported. Wider formats would require significant additional implementation 
work, not just selecting a larger precision.

Porting libbid to Java reuses established algorithms and test vectors, but does 
not automatically transfer the original implementation's maturity. Roughly 20K 
lines of newly ported arithmetic still create a substantial obligation for 
correctness, review, and long-term maintenance.

Java and Python already have mature arbitrary-precision decimal arithmetic. I 
would first evaluate a Spark-owned wrapper around BigDecimal for finite 
arithmetic, with explicit handling of special values  and error semantics. This 
could preserve wider precision and much less custom arithmetic code—a 
particularly important tradeoff for the main use case of decimal-floating point 
types: financial workloads.

Best,
Thomas

On Tue, Sep 22, 2026 at 4:17 PM Dongjoon Hyun 
<[email protected]<mailto:[email protected]>> wrote:
Thank you for the revised proposal and for the follow-up.

Let me start with my overall assessment, since I said the same thing
on the vote thread in August: this SPIP is still not mature enough to
go to a vote. The revision fixes one architectural point, but the
document still carries [OPEN] markers on its public API surface,
assumes a Parquet encoding that the Parquet community has explicitly
deferred, and leans on a config gate and a timeline whose precedent
does not support them. Details below.

On DB's specific question: yes, Revision 1 addresses the concern
behind my -1. The arithmetic path no longer depends on a native
library, and the fallback story is now a pure-Java, in-tree
implementation (Appendix E, item 6; SPARK-59111). That was the
blocking factor for me, and it is resolved at the SPIP level. One
follow-up on the same item: it still says native kernels "remain an
optional optimization under the same IEEE contract". I'd like the SPIP
to state plainly that the pure-Java path is the only arithmetic path
in v1, and that any native or accelerator path is out of scope and
would need its own SPIP. Otherwise the Java implementation risks being
demoted to a reference path that only the test suite exercises.

That said, I want to be precise about what exists today. The only code
on the table is PR #58410, which adds a ported BID arithmetic library
under a non-Spark package (org.bidfp): about 21K lines of Java, 10K
lines of tests, and 129K lines of Intel test vectors. By its own
description it "does not yet connect it to a user-facing Spark API or
execution path". There is still no DecFloatType, no parser or Catalyst
integration, no cast or coercion code, and no data source path. In
other words, the SPIP has been revised on paper, but the implementation
that would let us judge the pure-Java claim in practice has not been
provided. The PR itself has not been reviewed yet either. For that
reason I don't think the library should be merged as a standalone
module with no consumer in the tree. It should land together with the
first code that actually uses it, so that it is reviewed and exercised
as part of the feature rather than parked as dead code.

The SPIP-level items I'd want settled before a new vote are contracts
rather than implementation details:

1. Parquet. Appendix D assumes a BID little-endian FIXED_LEN_BYTE_ARRAY
   of 8/16 bytes. That is not where the parquet-dev thread is. As of
   the Sep 5 and Sep 21 messages, the points both sides agree on are:
   precision parameterized rather than fixed at 38, with narrower
   encodings allowed and a path to wider ones; coverage of the full
   finite decimal128 range; at minimum +/-Inf and a canonical NaN; and,
   explicitly, that the physical representation is to be selected only
   after public benchmarks and validation vectors. Canonicalization and
   sNaN/-0 handling are still open, and a joint proposal has not been
   written. I'd suggest the SPIP make Parquet persistence conditional
   on the Parquet logical type being adopted, and state that Spark will
   not ship a private encoding in the meantime. Russell raised the
   alignment question on the vote thread and I share it.

2. External types and Arrow. Appendix B (Java/Scala external type,
   JDBC) and the Arrow transport question are still marked [OPEN].
   These define what Row.get, Spark Connect, and Arrow-based collect in
   PySpark return, so they are public API. Ian's -0 asked for the Arrow
   half of this, and the revision does not answer it yet.

3. Gate and timeline. The SPIP cites the TIME type as the incubation
   pattern and bases the 9-12 month estimate on it. TIME is the worst
   possible example to lean on. Its SPIP passed on 2025-02-26. The type
   was developed without a gate for about nine months, and
   spark.sql.timeType.enabled was added on 2025-12-10, six days before
   the 4.1.0 release and during the RC phase, because support was
   incomplete (SPARK-54609). Nineteen months after the vote it is still
   off by default everywhere: in 4.1.0, in 4.2.0, in branch-4.3 whose
   RC1 was cut on 2026-08-31, in branch-4.x (4.4.0-SNAPSHOT), and on
   master, which is already Spark 5.0.0-SNAPSHOT. Coverage commits for
   TIME are still landing this month. So a config gate by itself is not
   a mitigation; it defers the risk across an entire major version
   line. DECFLOAT must not follow that precedent. We should not accept
   a second data type that lands piecemeal behind a flag with no
   defined end state, and I would treat repeating the TIME pattern as
   a reason to block rather than as an established practice to cite.
   Concretely, I'd ask the SPIP to define measurable exit criteria for
   the gate (function list, data sources, clients, plus the Parquet and
   Arrow contracts above), name who owns reaching them, and revisit the
   estimate, since DECFLOAT is strictly larger than TIME.

A few smaller design points in the current text should also be fixed:

- Coercion note (3) claims that lossy DECIMAL(35..38) -> DECFLOAT(34)
  widening is not implicit in ANSI mode, "consistent with ANSI
  coercion". Spark's ANSI mode already widens DECIMAL(38) + DOUBLE
  implicitly to DOUBLE, so the note describes a rule Spark does not
  have. The precedence also stops being a total order, so the result of
  DOUBLE + DECIMAL(38) + DECFLOAT(34) would depend on operand order.
- Appendix A declares DecFloatType as an AtomicType. Arithmetic and
  aggregates such as SUM and ABS accept NumericType, so it needs to be
  a FractionalType like DecimalType and DoubleType.
- Ordering is described as IEEE totalOrder, which distinguishes -0 from
  +0. Spark's DOUBLE treats -0.0 = 0.0 and NaN = NaN for equality,
  grouping, and joins, and sorts NaN after +Inf. DECFLOAT should follow
  those conventions, or equality and ORDER BY will disagree.

Thanks,
Dongjoon.

---------------------------------------------------------------------
To unsubscribe e-mail: 
[email protected]<mailto:[email protected]>



--
THOMAS KISSINGER
Staff Software Engineer

MOBILE  +49 174-2195270
EMAIL  [email protected]<mailto:[email protected]>

[https://www.snowflake.com/wp-content/themes/snowflake/img/[email protected]]
Snowflake Inc.
135 Constitution Drive
Menlo Park, CA 94025, USA

Reply via email to