Hi, Matt: Thanks for raising this.
1) Apache governance: I would +1 for putting this in an apache repo, for example a sub repo of datafusion project. Shawn has stated most of the reasons, so I don't want to repeat it again. 2) Where do tests live/how tests should be maintained? Initially I was thinking about putting sqllogictest in iceberg-rust, but after second thought I'm leaning towards to put it in the new repo for two reasons: 1. It would be easier for developer of the datafusion-iceberg integration to add tests 2. It would make the dependency graph and version release easier. Though the dependency is on crate level rather than repo level, the bi-direction dependency may make version management weird and difficult. The downside of this approach is that it's a little unfriendly for iceberg-rust developers, but I think it's less frequent for iceberg-rust developers to add sqllogictests compared with datafusion-iceberg developers. 3) Dropping DataFusion dependencies from pyiceberg-core binding I'm not quite familiar with this part, but it sounds reasonable to me. On Sat, Aug 22, 2026 at 6:50 AM Shawn Chang <[email protected]> wrote: > Hi Matt, > > Thanks for raising this! I generally agree that we can move the datafusion > integration to a separate repo if that helps more DataFusion experts to > work on the integrations > > On the three considerations: > 1) Apache governance: > I think this is my biggest concern so far. This could not only affect > contributors, but also whether the downstream users could continue to use > the integration. > Even if the code remains Apache-licensed, governance, release of artifacts, > and contribution policies can matter for adoption. > We should explore more about the option to keep it under an Apache-governed > repository. > > 2) Where do tests live/how tests should be maintained? > > I think this is closely related to the governance question. To me, the > broader question is: *how do we make it easy for people from both the > Iceberg Rust and DataFusion communities to maintain the integration?* > > If we move it to a non-Apache repository, I worry that it could eventually > look somewhat like the current Iceberg Java <> Trino integration: the > integration primarily lives on the Trino side and is therefore mostly > maintained by people who are already deeply involved in Trino. > > There is an important difference here, though. For Iceberg Java, Spark is > arguably the primary engine integration and has a large Iceberg contributor > base around it, so having the Trino integration maintained more > independently is relatively natural. In Iceberg Rust today, DataFusion has > a much more central role. It is by far the most mature engine integration > in the project, is used by our SQLLogicTest infrastructure, and many users > building on Iceberg Rust are also building on DataFusion. The overlap > between the two communities is therefore much larger. > > Because of that, I would prefer that extracting the integration does not > turn it into something that is effectively owned only by the DataFusion > side. Ideally, both Iceberg Rust contributors and DataFusion contributors > should be able to review changes, maintain compatibility, and evolve the > integration together. > > 3) Dropping DataFusion dependencies from pyiceberg-core binding: I think it > makes sense to make things simpler and have left a comment on the github > issue with more detailed thoughts: > https://github.com/apache/iceberg-rust/issues/3036 > > Best, > Shawn > > On Fri, Aug 21, 2026 at 11:23 AM Matt Butrovich <[email protected]> > wrote: > > > Hello all, > > > > > > I've never tried emailing two different project lists at once, but it was > > suggested that I do so to try to track the conversation across both > > communities. We'll see how this threads on the mailing lists. > > > > > > There has been a GitHub Discussion > > > https://github.com/apache/iceberg-rust/discussions/2992#discussioncomment-18082533 > , > > a GitHub Issue https://github.com/apache/iceberg-rust/issues/3029, and > > it's been a long topic of conversation in the past two weeks in both the > > DataFusion and Iceberg Rust Community Calls to discuss moving the > > DataFusion integration from Iceberg Rust to a separate repository. I will > > try to summarize some of the major points as I understand them, but the > > conversations are the ground truth and please feel free to correct me > here. > > > > > > The DataFusion TableProvider integration in Iceberg Rust makes DataFusion > > a dependency for Iceberg Rust. The integration exists for multiple > reasons: > > > > 1) an engine to execute Iceberg Rust's corpus of sqllogictest files > > > > 2) a TableProvider integration for DataFusion users to interact with > > Iceberg tables > > > > > > Some of the motivations to break out the integration: > > > > 1) There have been a number of issues and pull requests against the > > DataFusion TableProvider in Iceberg Rust as users want to add more > > features, and they often go stale. I don't believe there are many > > committers/PMC members familiar with or using the DataFusion integration. > > > > 2) Iceberg Rust would like to stay as engine-agnostic as possible. A > > recent DataFusion Ballista integration was declined for this reason > > https://github.com/apache/iceberg-rust/pull/2613. > > > > 3) Other projects that rely on both Iceberg Rust and DataFusion (e.g., > > Comet) are blocked by Iceberg Rust upgrading its DataFusion and Arrow > > dependencies before they can upgrade. > > > > > > Some considerations for both communities: > > > > 1) Where would this DataFusion TableProvider live? Most specifically, > > would it be an Apache-governed project? As folks like @andygrove point > out, > > this can affect whether some community members could contribute to it. > > There is a datafusion-contrib org for DataFusion-related projects to have > > visibility but no Apache governance, but there may be options to put it > > under an Apache repository. > > > > 2) How would Iceberg Rust continue to run sqllogictests for regression > > testing? Does this live in a different repository that depends on this > new > > Iceberg Rust TableProvider crate? Would we be able to test Pull Requests > on > > Iceberg Rust with sqllogictests? > > > > 3) @kevinqliu is familiar with the Python bindings in Iceberg Rust, and I > > believe that DataFusion dependency might be removed as well, but that is > > not the core focus of this conversation. He has also proposed removing > that > > and using the DataFusion Python bindings directly. > > > > > > I'm sure I'm forgetting things, but this email is long enough. > > > > > > Thanks everyone for the discussion thus far, and looking forward to more > > input on this. > > > > > > -Matt > > > > > > > > >
