Thanks Tim. I left review comments on the RFC in PR #612.They are mostly
questions rather than blockers, and I am +1 on the direction.

On Thu, Aug 13, 2026 05:16 AM, Tim Brown <[email protected]> wrote:

> I will revive the changes in PR #618. I plan on breaking down the full set
> of the changes to first get the test framework setup and merged, then I
> will break out the individual modules.
> It would be good to have some initial review on the RFC before we go much
> further to make sure things are aligned
> https://github.com/apache/incubator-xtable/pull/612
>
> On Mon, Aug 10, 2026 at 7:34 PM Vinish <[email protected]>
> wrote:
>
> > Hi Tim,
> >
> > I am withdrawing the xtable-java-runtime proposal. A new module needs
> > application-specific code, and this one does not have any.
> >
> > Do you want to revive PR #618, or should someone pick it up against the
> > current main? It predates Paimon, Parquet and Kernel.
> >
> > On your point about Delta Kernel, I have filed #886 for the
> > ITConversionController gap.
> > https://github.com/apache/incubator-xtable/issues/886.
> > We will pick it up regardless of how the packaging question lands, since
> it
> > is the precondition for making Kernel the default Delta path.
> >
> > Thanks,
> > Vinish
> >
> > On Sun, Jul 19, 2026 08:10 AM, Tim Brown <[email protected]> wrote:
> >
> > > If the goal is to make it easier for java applications, I think we
> should
> > > just go with the proposal in RFC-2. I have a now outdated
> implementation
> > > [1] of this including a bundle validation flow. The proposed jar is
> > > excluding Paimon for some reason and I think we should aim to support
> all
> > > the table formats that are supported. Bundling all of this will just
> lead
> > > us back to an overly large jar so it is key to keep things modular.
> > >
> > > Generally when people are building java applications, they are building
> > > bundled jars when they want to run these applications so it doesn't
> seem
> > > like a bundled jar is really the requirement here since they will be
> > > bundling the required dependencies declared by our jar's manifest. The
> > need
> > > here is to make sure the manifest is as lean as possible so someone
> that
> > > just wants delta-kernel and Iceberg will not pull in Hudi, Paimon, or
> any
> > > spark dependency.
> > >
> > > If there is a secondary concern that the user will already have their
> own
> > > version of some of the required dependencies that may conflict with
> what
> > we
> > > declare, then it makes sense to also publish our own bundled jars with
> > > relocated dependencies for these lean modules.
> > >
> > > The only reason I can see for adding yet another module to the project
> is
> > > if there is some application specific code that needs to be added. This
> > is
> > > the case for the spark proposal since it will include pre-built
> solutions
> > > for how to integrate XTable into existing writers.
> > >
> > >
> > > 1.
> > >
> > >
> >
> https://github.com/apache/incubator-xtable/pull/618/changes#diff-b818b66a60f0848d9b846b5f966ca959be112f27a12694fa9f79002d9ee9e07e
> > >
> > > -Tim
> > >
> > > On Fri, Jul 17, 2026 at 12:58 PM Vaibhav Kumar <[email protected]
> >
> > > wrote:
> > >
> > > > Hi Tim,
> > > >
> > > >
> > > >
> > > > Thanks for the response and sharing the RFC.
> > > >
> > > > You're right that we need to be clear on the goal. Let me be precise
> as
> > > per
> > > > our discussion in the last sync call : the goal is not to replace or
> > > patch
> > > > the utilities jar. It is to allow XTable to be used as a drop-in
> > library
> > > in
> > > > JVM applications that do not run Spark at all — Flink jobs,
> > Trino/Presto
> > > > plugins, plain Java  services. The utility jar cannot serve that goal
> > > even
> > > > if we slimmed it down, because it is  designed as a CLI tool and
> > bundles
> > > > Spark by design.
> > > >
> > > >
> > > >   This actually lines up directly with the pain points your RFC-2 [1]
> > > > identifies:
> > > >
> > > >
> > > >   ▎ "When running the prebuilt utilities bundle jar, the user will
> > bring
> > > in
> > > > all the required
> > > >   ▎ dependencies which currently include three table formats, spark,
> > and
> > > > hadoop. This results in a
> > > >   ▎ very large jar containing more than the user really needs."
> > > >
> > > >
> > > >
> > > >   ▎ "When building your own jar, you can run into issues with
> conflicts
> > > in
> > > > versions for the
> > > >   ▎ specific table formats you want to use since the user will often
> > > > already have an implementation
> > > >   ▎ of at least one format on their classpath."
> > > >
> > > > An xtable-java-runtime addresses both pain points directly for the
> > > > non-Spark user segment. It  would be a shaded jar with relocated
> > > > dependencies — the same pattern used in xtable-spark-runtime(#843) —
> > > > bundling Delta Kernel, Hudi Java client, and Iceberg core with no
> > Spark.
> > > > Because dependencies are relocated inside the jar, version conflicts
> on
> > > the
> > > > user's classpath are avoided. This is exactly the per-format shaded
> jar
> > > > approach your RFC-2 proposes.
> > > >
> > > >
> > > > To your point on jar structure: it would be a shaded fat jar with
> > > relocated
> > > > dependencies. Users would consume it via --jars or as a Maven
> > dependency,
> > > > and the dependency-reduced POM would expose no transitive engine
> > > > dependencies.
> > > >
> > > >
> > > > On Delta Kernel readiness — you are right that it is not yet
> integrated
> > > > into ITConversionController as a first-class path alongside the
> > > Spark-based
> > > > Delta tests. The dedicated test suite (ITDeltaKernelConversionSource,
> > > > TestDeltaKernelReadWriteIntegration,  TestDeltaKernelSync) covers the
> > > > Kernel path in depth, but I agree the shared harness is the right
> bar.
> > > > Would it make sense to track the ITConversionController integration
> as
> > a
> > > > prerequisite for the java-runtime publish, so we can keep the
> > discussion
> > > > moving forward while that work gets done in parallel?
> > > >
> > > >
> > > > To summarise — the xtable-java-runtime is not a workaround for the
> > module
> > > > restructuring RFC. It is one concrete deliverable that RFC-2 leads
> to,
> > > > scoped narrowly to a packaging module so it does not depend on the
> full
> > > > restructure being complete first.
> > > >
> > > >
> > > > Happy to open a tracking issue for the ITConversionController gap and
> > put
> > > > together a more  detailed design doc if the community agrees with the
> > > > direction.
> > > >
> > > > Thanks,
> > > >
> > > > Vaibhav
> > > >
> > > >
> > > >
> > > > On Tue, Jul 14, 2026 at 5:53 PM Tim Brown <[email protected]>
> > > wrote:
> > > >
> > > > > I think that we need to figure out what the goal is before
> proposing
> > > the
> > > > > structure. If the goal is to allow a standalone java jar to run
> > > > > conversions, I would suggest we simply revisit the utilities jar or
> > the
> > > > web
> > > > > service for this.
> > > > >
> > > > > The main issue with the utilities jar was that it was bundling all
> > > sorts
> > > > of
> > > > > things, including spark. We have a similar issue with the layout of
> > the
> > > > > project in general where it is hard to pick and choose the lean set
> > of
> > > > > dependencies that you want for your particular use case. That was
> the
> > > > > reason for the RFC for module restructuring [1].
> > > > >
> > > > > Before starting any work on this there are a few questions that
> need
> > to
> > > > be
> > > > > answered. What will the xtable-java-runtime jar look like? Will it
> > be a
> > > > > bundled jar? Will it have relocated dependencies or will it require
> > > users
> > > > > to provide the implementation of the table formats?
> > > > >
> > > > > Before declaring the Delta Kernel ready for use, I would like to
> see
> > it
> > > > > integrated into our integration test suite in
> ITConversionController
> > > [2].
> > > > > It is also worthwhile to look back at the recent issues reported
> for
> > > > Delta
> > > > > conversion and ensure they are also fixed in the Delta Kernel path.
> > > > >
> > > > > 1. https://github.com/apache/incubator-xtable/pull/612
> > > > > 2.
> > > > >
> > > > >
> > > >
> > >
> >
> https://github.com/apache/incubator-xtable/blob/main/xtable-core/src/test/java/org/apache/xtable/ITConversionController.java
> > > > >
> > > > > On Mon, Jul 13, 2026 at 10:31 PM Vaibhav Kumar <
> > [email protected]
> > > >
> > > > > wrote:
> > > > >
> > > > > > Hey Vinish
> > > > > >
> > > > > > +1 for the idea, I can take this up.
> > > > > >
> > > > > > On Tue, 14 Jul 2026 at 6:39 AM, Vinish Reddy Pannala <
> > > > > > [email protected]> wrote:
> > > > > >
> > > > > > > Hi all,
> > > > > > >
> > > > > > > I'd like to open a discussion on publishing an
> > xtable-java-runtime
> > > -
> > > > a
> > > > > > > lightweight, Spark-free runtime for running XTable metadata
> sync
> > -
> > > > now
> > > > > > that
> > > > > > > Delta Kernel support has landed on main.
> > > > > > >
> > > > > > > Motivation
> > > > > > >
> > > > > > > Until now, any XTable sync involving Delta effectively
> required a
> > > > > > > SparkSession, since the Delta source/target went through
> > > > delta-spark's
> > > > > > > DeltaLog APIs. The Hudi and Iceberg paths are already pure
> Java,
> > so
> > > > > Delta
> > > > > > > was the one format anchoring us to Spark. Today the runtime
> > options
> > > > we
> > > > > > ship
> > > > > > > (the xtable-utilities CLI and xtable-spark-runtime, #843) both
> > > carry
> > > > > > Spark.
> > > > > > >
> > > > > > > With Delta Kernel PR's now merged, Delta can be read and
> written
> > > > > without
> > > > > > > Spark. That removes the last hard Spark dependency in the sync
> > > path,
> > > > > > which
> > > > > > > means we can publish a genuinely Spark-free, pure-Java runtime
> > jar
> > > > > > covering
> > > > > > > Hudi, Iceberg, and Delta (via Kernel).
> > > > > > >
> > > > > > > Why it matters - non-Spark engines
> > > > > > >
> > > > > > > A Java-only runtime makes XTable embeddable anywhere on the JVM
> > > > without
> > > > > > > dragging in Spark:
> > > > > > >
> > > > > > > - Flink jobs syncing metadata in-line
> > > > > > > - Trino / Presto plugins and connectors
> > > > > > > - Plain Java services, functions, or orchestration steps
> > > > > > >
> > > > > > > This drastically shrinks the footprint and the
> > dependency-conflict
> > > > > > surface
> > > > > > > versus a Spark bundle, and opens XTable to the large set of
> users
> > > who
> > > > > > > aren't on Spark.
> > > > > > >
> > > > > > > Context
> > > > > > >
> > > > > > > We discussed this during the community sync this morning and
> > agreed
> > > > > it's
> > > > > > > worth pursuing - the notes capture that with Delta Kernel
> merged,
> > > an
> > > > > > > xtable-java-runtime jar can be published with Delta Kernel as
> > > > > > > source/target. Full notes:
> > > > > > >
> > > > > > >
> > > > > >
> > > > >
> > > >
> > >
> >
> https://docs.google.com/document/d/1mSthtQBVDDzi9bLn9sWDsPaJLJHCDoK_MxDsSphhlos/edit?usp=sharing
> > > > > > >
> > > > > > > Questions for the list
> > > > > > >
> > > > > > > - Do we agree an xtable-java-runtime is worth publishing in the
> > > next
> > > > > > > release?
> > > > > > > - Any concerns about Delta Kernel feature parity for a
> > Kernel-only
> > > > > > runtime
> > > > > > > (e.g. deletion vectors, issue #713)?
> > > > > > >
> > > > > > > Looking forward to your thoughts.
> > > > > > >
> > > > > > > Thanks,
> > > > > > > Vinish
> > > > > > >
> > > > > >
> > > > >
> > > >
> > >
> >
>

Reply via email to