I will revive the changes in PR #618. I plan on breaking down the full set
of the changes to first get the test framework setup and merged, then I
will break out the individual modules.
It would be good to have some initial review on the RFC before we go much
further to make sure things are aligned
https://github.com/apache/incubator-xtable/pull/612

On Mon, Aug 10, 2026 at 7:34 PM Vinish <[email protected]> wrote:

> Hi Tim,
>
> I am withdrawing the xtable-java-runtime proposal. A new module needs
> application-specific code, and this one does not have any.
>
> Do you want to revive PR #618, or should someone pick it up against the
> current main? It predates Paimon, Parquet and Kernel.
>
> On your point about Delta Kernel, I have filed #886 for the
> ITConversionController gap.
> https://github.com/apache/incubator-xtable/issues/886.
> We will pick it up regardless of how the packaging question lands, since it
> is the precondition for making Kernel the default Delta path.
>
> Thanks,
> Vinish
>
> On Sun, Jul 19, 2026 08:10 AM, Tim Brown <[email protected]> wrote:
>
> > If the goal is to make it easier for java applications, I think we should
> > just go with the proposal in RFC-2. I have a now outdated implementation
> > [1] of this including a bundle validation flow. The proposed jar is
> > excluding Paimon for some reason and I think we should aim to support all
> > the table formats that are supported. Bundling all of this will just lead
> > us back to an overly large jar so it is key to keep things modular.
> >
> > Generally when people are building java applications, they are building
> > bundled jars when they want to run these applications so it doesn't seem
> > like a bundled jar is really the requirement here since they will be
> > bundling the required dependencies declared by our jar's manifest. The
> need
> > here is to make sure the manifest is as lean as possible so someone that
> > just wants delta-kernel and Iceberg will not pull in Hudi, Paimon, or any
> > spark dependency.
> >
> > If there is a secondary concern that the user will already have their own
> > version of some of the required dependencies that may conflict with what
> we
> > declare, then it makes sense to also publish our own bundled jars with
> > relocated dependencies for these lean modules.
> >
> > The only reason I can see for adding yet another module to the project is
> > if there is some application specific code that needs to be added. This
> is
> > the case for the spark proposal since it will include pre-built solutions
> > for how to integrate XTable into existing writers.
> >
> >
> > 1.
> >
> >
> https://github.com/apache/incubator-xtable/pull/618/changes#diff-b818b66a60f0848d9b846b5f966ca959be112f27a12694fa9f79002d9ee9e07e
> >
> > -Tim
> >
> > On Fri, Jul 17, 2026 at 12:58 PM Vaibhav Kumar <[email protected]>
> > wrote:
> >
> > > Hi Tim,
> > >
> > >
> > >
> > > Thanks for the response and sharing the RFC.
> > >
> > > You're right that we need to be clear on the goal. Let me be precise as
> > per
> > > our discussion in the last sync call : the goal is not to replace or
> > patch
> > > the utilities jar. It is to allow XTable to be used as a drop-in
> library
> > in
> > > JVM applications that do not run Spark at all — Flink jobs,
> Trino/Presto
> > > plugins, plain Java  services. The utility jar cannot serve that goal
> > even
> > > if we slimmed it down, because it is  designed as a CLI tool and
> bundles
> > > Spark by design.
> > >
> > >
> > >   This actually lines up directly with the pain points your RFC-2 [1]
> > > identifies:
> > >
> > >
> > >   ▎ "When running the prebuilt utilities bundle jar, the user will
> bring
> > in
> > > all the required
> > >   ▎ dependencies which currently include three table formats, spark,
> and
> > > hadoop. This results in a
> > >   ▎ very large jar containing more than the user really needs."
> > >
> > >
> > >
> > >   ▎ "When building your own jar, you can run into issues with conflicts
> > in
> > > versions for the
> > >   ▎ specific table formats you want to use since the user will often
> > > already have an implementation
> > >   ▎ of at least one format on their classpath."
> > >
> > > An xtable-java-runtime addresses both pain points directly for the
> > > non-Spark user segment. It  would be a shaded jar with relocated
> > > dependencies — the same pattern used in xtable-spark-runtime(#843) —
> > > bundling Delta Kernel, Hudi Java client, and Iceberg core with no
> Spark.
> > > Because dependencies are relocated inside the jar, version conflicts on
> > the
> > > user's classpath are avoided. This is exactly the per-format shaded jar
> > > approach your RFC-2 proposes.
> > >
> > >
> > > To your point on jar structure: it would be a shaded fat jar with
> > relocated
> > > dependencies. Users would consume it via --jars or as a Maven
> dependency,
> > > and the dependency-reduced POM would expose no transitive engine
> > > dependencies.
> > >
> > >
> > > On Delta Kernel readiness — you are right that it is not yet integrated
> > > into ITConversionController as a first-class path alongside the
> > Spark-based
> > > Delta tests. The dedicated test suite (ITDeltaKernelConversionSource,
> > > TestDeltaKernelReadWriteIntegration,  TestDeltaKernelSync) covers the
> > > Kernel path in depth, but I agree the shared harness is the right bar.
> > > Would it make sense to track the ITConversionController integration as
> a
> > > prerequisite for the java-runtime publish, so we can keep the
> discussion
> > > moving forward while that work gets done in parallel?
> > >
> > >
> > > To summarise — the xtable-java-runtime is not a workaround for the
> module
> > > restructuring RFC. It is one concrete deliverable that RFC-2 leads to,
> > > scoped narrowly to a packaging module so it does not depend on the full
> > > restructure being complete first.
> > >
> > >
> > > Happy to open a tracking issue for the ITConversionController gap and
> put
> > > together a more  detailed design doc if the community agrees with the
> > > direction.
> > >
> > > Thanks,
> > >
> > > Vaibhav
> > >
> > >
> > >
> > > On Tue, Jul 14, 2026 at 5:53 PM Tim Brown <[email protected]>
> > wrote:
> > >
> > > > I think that we need to figure out what the goal is before proposing
> > the
> > > > structure. If the goal is to allow a standalone java jar to run
> > > > conversions, I would suggest we simply revisit the utilities jar or
> the
> > > web
> > > > service for this.
> > > >
> > > > The main issue with the utilities jar was that it was bundling all
> > sorts
> > > of
> > > > things, including spark. We have a similar issue with the layout of
> the
> > > > project in general where it is hard to pick and choose the lean set
> of
> > > > dependencies that you want for your particular use case. That was the
> > > > reason for the RFC for module restructuring [1].
> > > >
> > > > Before starting any work on this there are a few questions that need
> to
> > > be
> > > > answered. What will the xtable-java-runtime jar look like? Will it
> be a
> > > > bundled jar? Will it have relocated dependencies or will it require
> > users
> > > > to provide the implementation of the table formats?
> > > >
> > > > Before declaring the Delta Kernel ready for use, I would like to see
> it
> > > > integrated into our integration test suite in ITConversionController
> > [2].
> > > > It is also worthwhile to look back at the recent issues reported for
> > > Delta
> > > > conversion and ensure they are also fixed in the Delta Kernel path.
> > > >
> > > > 1. https://github.com/apache/incubator-xtable/pull/612
> > > > 2.
> > > >
> > > >
> > >
> >
> https://github.com/apache/incubator-xtable/blob/main/xtable-core/src/test/java/org/apache/xtable/ITConversionController.java
> > > >
> > > > On Mon, Jul 13, 2026 at 10:31 PM Vaibhav Kumar <
> [email protected]
> > >
> > > > wrote:
> > > >
> > > > > Hey Vinish
> > > > >
> > > > > +1 for the idea, I can take this up.
> > > > >
> > > > > On Tue, 14 Jul 2026 at 6:39 AM, Vinish Reddy Pannala <
> > > > > [email protected]> wrote:
> > > > >
> > > > > > Hi all,
> > > > > >
> > > > > > I'd like to open a discussion on publishing an
> xtable-java-runtime
> > -
> > > a
> > > > > > lightweight, Spark-free runtime for running XTable metadata sync
> -
> > > now
> > > > > that
> > > > > > Delta Kernel support has landed on main.
> > > > > >
> > > > > > Motivation
> > > > > >
> > > > > > Until now, any XTable sync involving Delta effectively required a
> > > > > > SparkSession, since the Delta source/target went through
> > > delta-spark's
> > > > > > DeltaLog APIs. The Hudi and Iceberg paths are already pure Java,
> so
> > > > Delta
> > > > > > was the one format anchoring us to Spark. Today the runtime
> options
> > > we
> > > > > ship
> > > > > > (the xtable-utilities CLI and xtable-spark-runtime, #843) both
> > carry
> > > > > Spark.
> > > > > >
> > > > > > With Delta Kernel PR's now merged, Delta can be read and written
> > > > without
> > > > > > Spark. That removes the last hard Spark dependency in the sync
> > path,
> > > > > which
> > > > > > means we can publish a genuinely Spark-free, pure-Java runtime
> jar
> > > > > covering
> > > > > > Hudi, Iceberg, and Delta (via Kernel).
> > > > > >
> > > > > > Why it matters - non-Spark engines
> > > > > >
> > > > > > A Java-only runtime makes XTable embeddable anywhere on the JVM
> > > without
> > > > > > dragging in Spark:
> > > > > >
> > > > > > - Flink jobs syncing metadata in-line
> > > > > > - Trino / Presto plugins and connectors
> > > > > > - Plain Java services, functions, or orchestration steps
> > > > > >
> > > > > > This drastically shrinks the footprint and the
> dependency-conflict
> > > > > surface
> > > > > > versus a Spark bundle, and opens XTable to the large set of users
> > who
> > > > > > aren't on Spark.
> > > > > >
> > > > > > Context
> > > > > >
> > > > > > We discussed this during the community sync this morning and
> agreed
> > > > it's
> > > > > > worth pursuing - the notes capture that with Delta Kernel merged,
> > an
> > > > > > xtable-java-runtime jar can be published with Delta Kernel as
> > > > > > source/target. Full notes:
> > > > > >
> > > > > >
> > > > >
> > > >
> > >
> >
> https://docs.google.com/document/d/1mSthtQBVDDzi9bLn9sWDsPaJLJHCDoK_MxDsSphhlos/edit?usp=sharing
> > > > > >
> > > > > > Questions for the list
> > > > > >
> > > > > > - Do we agree an xtable-java-runtime is worth publishing in the
> > next
> > > > > > release?
> > > > > > - Any concerns about Delta Kernel feature parity for a
> Kernel-only
> > > > > runtime
> > > > > > (e.g. deletion vectors, issue #713)?
> > > > > >
> > > > > > Looking forward to your thoughts.
> > > > > >
> > > > > > Thanks,
> > > > > > Vinish
> > > > > >
> > > > >
> > > >
> > >
> >
>

Reply via email to