I will revive the changes in PR #618. I plan on breaking down the full set of the changes to first get the test framework setup and merged, then I will break out the individual modules. It would be good to have some initial review on the RFC before we go much further to make sure things are aligned https://github.com/apache/incubator-xtable/pull/612
On Mon, Aug 10, 2026 at 7:34 PM Vinish <[email protected]> wrote: > Hi Tim, > > I am withdrawing the xtable-java-runtime proposal. A new module needs > application-specific code, and this one does not have any. > > Do you want to revive PR #618, or should someone pick it up against the > current main? It predates Paimon, Parquet and Kernel. > > On your point about Delta Kernel, I have filed #886 for the > ITConversionController gap. > https://github.com/apache/incubator-xtable/issues/886. > We will pick it up regardless of how the packaging question lands, since it > is the precondition for making Kernel the default Delta path. > > Thanks, > Vinish > > On Sun, Jul 19, 2026 08:10 AM, Tim Brown <[email protected]> wrote: > > > If the goal is to make it easier for java applications, I think we should > > just go with the proposal in RFC-2. I have a now outdated implementation > > [1] of this including a bundle validation flow. The proposed jar is > > excluding Paimon for some reason and I think we should aim to support all > > the table formats that are supported. Bundling all of this will just lead > > us back to an overly large jar so it is key to keep things modular. > > > > Generally when people are building java applications, they are building > > bundled jars when they want to run these applications so it doesn't seem > > like a bundled jar is really the requirement here since they will be > > bundling the required dependencies declared by our jar's manifest. The > need > > here is to make sure the manifest is as lean as possible so someone that > > just wants delta-kernel and Iceberg will not pull in Hudi, Paimon, or any > > spark dependency. > > > > If there is a secondary concern that the user will already have their own > > version of some of the required dependencies that may conflict with what > we > > declare, then it makes sense to also publish our own bundled jars with > > relocated dependencies for these lean modules. > > > > The only reason I can see for adding yet another module to the project is > > if there is some application specific code that needs to be added. This > is > > the case for the spark proposal since it will include pre-built solutions > > for how to integrate XTable into existing writers. > > > > > > 1. > > > > > https://github.com/apache/incubator-xtable/pull/618/changes#diff-b818b66a60f0848d9b846b5f966ca959be112f27a12694fa9f79002d9ee9e07e > > > > -Tim > > > > On Fri, Jul 17, 2026 at 12:58 PM Vaibhav Kumar <[email protected]> > > wrote: > > > > > Hi Tim, > > > > > > > > > > > > Thanks for the response and sharing the RFC. > > > > > > You're right that we need to be clear on the goal. Let me be precise as > > per > > > our discussion in the last sync call : the goal is not to replace or > > patch > > > the utilities jar. It is to allow XTable to be used as a drop-in > library > > in > > > JVM applications that do not run Spark at all — Flink jobs, > Trino/Presto > > > plugins, plain Java services. The utility jar cannot serve that goal > > even > > > if we slimmed it down, because it is designed as a CLI tool and > bundles > > > Spark by design. > > > > > > > > > This actually lines up directly with the pain points your RFC-2 [1] > > > identifies: > > > > > > > > > ▎ "When running the prebuilt utilities bundle jar, the user will > bring > > in > > > all the required > > > ▎ dependencies which currently include three table formats, spark, > and > > > hadoop. This results in a > > > ▎ very large jar containing more than the user really needs." > > > > > > > > > > > > ▎ "When building your own jar, you can run into issues with conflicts > > in > > > versions for the > > > ▎ specific table formats you want to use since the user will often > > > already have an implementation > > > ▎ of at least one format on their classpath." > > > > > > An xtable-java-runtime addresses both pain points directly for the > > > non-Spark user segment. It would be a shaded jar with relocated > > > dependencies — the same pattern used in xtable-spark-runtime(#843) — > > > bundling Delta Kernel, Hudi Java client, and Iceberg core with no > Spark. > > > Because dependencies are relocated inside the jar, version conflicts on > > the > > > user's classpath are avoided. This is exactly the per-format shaded jar > > > approach your RFC-2 proposes. > > > > > > > > > To your point on jar structure: it would be a shaded fat jar with > > relocated > > > dependencies. Users would consume it via --jars or as a Maven > dependency, > > > and the dependency-reduced POM would expose no transitive engine > > > dependencies. > > > > > > > > > On Delta Kernel readiness — you are right that it is not yet integrated > > > into ITConversionController as a first-class path alongside the > > Spark-based > > > Delta tests. The dedicated test suite (ITDeltaKernelConversionSource, > > > TestDeltaKernelReadWriteIntegration, TestDeltaKernelSync) covers the > > > Kernel path in depth, but I agree the shared harness is the right bar. > > > Would it make sense to track the ITConversionController integration as > a > > > prerequisite for the java-runtime publish, so we can keep the > discussion > > > moving forward while that work gets done in parallel? > > > > > > > > > To summarise — the xtable-java-runtime is not a workaround for the > module > > > restructuring RFC. It is one concrete deliverable that RFC-2 leads to, > > > scoped narrowly to a packaging module so it does not depend on the full > > > restructure being complete first. > > > > > > > > > Happy to open a tracking issue for the ITConversionController gap and > put > > > together a more detailed design doc if the community agrees with the > > > direction. > > > > > > Thanks, > > > > > > Vaibhav > > > > > > > > > > > > On Tue, Jul 14, 2026 at 5:53 PM Tim Brown <[email protected]> > > wrote: > > > > > > > I think that we need to figure out what the goal is before proposing > > the > > > > structure. If the goal is to allow a standalone java jar to run > > > > conversions, I would suggest we simply revisit the utilities jar or > the > > > web > > > > service for this. > > > > > > > > The main issue with the utilities jar was that it was bundling all > > sorts > > > of > > > > things, including spark. We have a similar issue with the layout of > the > > > > project in general where it is hard to pick and choose the lean set > of > > > > dependencies that you want for your particular use case. That was the > > > > reason for the RFC for module restructuring [1]. > > > > > > > > Before starting any work on this there are a few questions that need > to > > > be > > > > answered. What will the xtable-java-runtime jar look like? Will it > be a > > > > bundled jar? Will it have relocated dependencies or will it require > > users > > > > to provide the implementation of the table formats? > > > > > > > > Before declaring the Delta Kernel ready for use, I would like to see > it > > > > integrated into our integration test suite in ITConversionController > > [2]. > > > > It is also worthwhile to look back at the recent issues reported for > > > Delta > > > > conversion and ensure they are also fixed in the Delta Kernel path. > > > > > > > > 1. https://github.com/apache/incubator-xtable/pull/612 > > > > 2. > > > > > > > > > > > > > > https://github.com/apache/incubator-xtable/blob/main/xtable-core/src/test/java/org/apache/xtable/ITConversionController.java > > > > > > > > On Mon, Jul 13, 2026 at 10:31 PM Vaibhav Kumar < > [email protected] > > > > > > > wrote: > > > > > > > > > Hey Vinish > > > > > > > > > > +1 for the idea, I can take this up. > > > > > > > > > > On Tue, 14 Jul 2026 at 6:39 AM, Vinish Reddy Pannala < > > > > > [email protected]> wrote: > > > > > > > > > > > Hi all, > > > > > > > > > > > > I'd like to open a discussion on publishing an > xtable-java-runtime > > - > > > a > > > > > > lightweight, Spark-free runtime for running XTable metadata sync > - > > > now > > > > > that > > > > > > Delta Kernel support has landed on main. > > > > > > > > > > > > Motivation > > > > > > > > > > > > Until now, any XTable sync involving Delta effectively required a > > > > > > SparkSession, since the Delta source/target went through > > > delta-spark's > > > > > > DeltaLog APIs. The Hudi and Iceberg paths are already pure Java, > so > > > > Delta > > > > > > was the one format anchoring us to Spark. Today the runtime > options > > > we > > > > > ship > > > > > > (the xtable-utilities CLI and xtable-spark-runtime, #843) both > > carry > > > > > Spark. > > > > > > > > > > > > With Delta Kernel PR's now merged, Delta can be read and written > > > > without > > > > > > Spark. That removes the last hard Spark dependency in the sync > > path, > > > > > which > > > > > > means we can publish a genuinely Spark-free, pure-Java runtime > jar > > > > > covering > > > > > > Hudi, Iceberg, and Delta (via Kernel). > > > > > > > > > > > > Why it matters - non-Spark engines > > > > > > > > > > > > A Java-only runtime makes XTable embeddable anywhere on the JVM > > > without > > > > > > dragging in Spark: > > > > > > > > > > > > - Flink jobs syncing metadata in-line > > > > > > - Trino / Presto plugins and connectors > > > > > > - Plain Java services, functions, or orchestration steps > > > > > > > > > > > > This drastically shrinks the footprint and the > dependency-conflict > > > > > surface > > > > > > versus a Spark bundle, and opens XTable to the large set of users > > who > > > > > > aren't on Spark. > > > > > > > > > > > > Context > > > > > > > > > > > > We discussed this during the community sync this morning and > agreed > > > > it's > > > > > > worth pursuing - the notes capture that with Delta Kernel merged, > > an > > > > > > xtable-java-runtime jar can be published with Delta Kernel as > > > > > > source/target. Full notes: > > > > > > > > > > > > > > > > > > > > > > > > > > > https://docs.google.com/document/d/1mSthtQBVDDzi9bLn9sWDsPaJLJHCDoK_MxDsSphhlos/edit?usp=sharing > > > > > > > > > > > > Questions for the list > > > > > > > > > > > > - Do we agree an xtable-java-runtime is worth publishing in the > > next > > > > > > release? > > > > > > - Any concerns about Delta Kernel feature parity for a > Kernel-only > > > > > runtime > > > > > > (e.g. deletion vectors, issue #713)? > > > > > > > > > > > > Looking forward to your thoughts. > > > > > > > > > > > > Thanks, > > > > > > Vinish > > > > > > > > > > > > > > > > > > > > >
