I think this is a great idea as well -- thank you for bringing it up and for the great discussions so far.
While at VLDB this past week, I spoke to at least three people from companies adding Apache Iceberg support to their products. All of them had to fork iceberg-rust for one reason or another. I think we have a huge need to improve our ability to work together and accelerate everyone's efforts. Giving the DataFusion integration access to more expert maintainers with bandwidth I think will help everyone. There seems to be consensus that we should move the integration code to DataFusion governance, and that it should remain in the ASF, though some open technical questions remain. If that is the case, I propose the following specific process: 1. Create the new gitub repository in Apache for the code (e.g. apache/datafusion-iceberg) 2. Create a PR in the new repo with the proposed code 3. Hold a formal vote on the iceberg dev list to move the integration code to DataFusion 4. Hold a formal vote on the DataFusion dev list to accept the new code I am happy to help with the logistics (e.g. ASF INFRA ticket to create the new repo, votes, etc) but I am not expert enough to create the proposed PR. Please let me know your thoughts, Andrew (PMC Chair of DataFusion) On Tue, Sep 1, 2026 at 9:48 PM Renjie Liu <[email protected]> wrote: > To add more background about the relationship between comet and datafusion > for those who are not familiar with them. > > Apache datafusion is a popular extensible compute engine written in rust. > Apache comet is an apache spark accelerator builton on apache datafusion, > and also a subproject of apache datafusion. > > datafusion-iceberg is an apache datafusion extension built on iceberg-rust, > and comet's iceberg support is built on it. > > On Tue, Sep 1, 2026 at 9:39 PM Andy Grove <[email protected]> wrote: > > > +1 for moving to DataFusion PMC. The iceberg-datafusion integration is > > very important for Comet. > > > > On Mon, Aug 31, 2026 at 11:42 PM Xuanwo <[email protected]> wrote: > > > > > > TBH, I also support moving to the DataFusion PMC. > > > > > > - DataFusion is the largest dependency in iceberg-datafusion. > > > - The largest downstream user of iceberg-datafusion is Comet, which > > shares many of the same PMC members as DataFusion. > > > > > > It feels natural to be part of the DataFusion PMC. As long as DF PMC is > > willing to accept this project, it LGTM. > > > > > > On Tue, Sep 1, 2026, at 12:03, Renjie Liu wrote: > > > > > > I don't think the testing should be a blocker of moving > > iceberg-datafusion out of iceberg-rust repo. From what I learn, most of > the > > sqllogictests are in pr of modifying iceberg-datafusion integration, > there > > are only few cases where we rely on sqllogictests to verify features. > > > > > > > How would that sound for the `iceberg-rust` community to grant a > > couple of Apache DataFusion PMCs committer access to the repository, > > limited to the `iceberg-datafusion` crate via a CODEOWNERS file, so that > > they can help push reviews and PRs forward independently? > > > > > > I'm not sure if this is feasible, but moving the iceberg-datafusion > > crate to apache datafusion project sounds a more reasonable approach to > me. > > It's still governed by Apache, and most of the code is related to > > DataFusion, so I think the DataFusion community is in a better position > to > > define the vision and design for it. > > > > > > On Mon, Aug 31, 2026 at 10:45 PM Gabriel Musat <[email protected]> > > wrote: > > > > > > Hi everyone, > > > > > > Based on these two facts: > > > - There's a current reviewer bandwidth problem that hurts development > > velocity of the `iceberg-datafusion` crate currently hosted under > > `apache/iceberg-rust`. > > > - There is value in maintaining the `iceberg-datafusion` crate inside > > `apache/iceberg-rust` for testing, design and governance. > > > > > > How would that sound for the `iceberg-rust` community to grant a couple > > of Apache DataFusion PMCs committer access to the repository, limited to > > the `iceberg-datafusion` crate via a CODEOWNERS file, so that they can > help > > push reviews and PRs forward independently? > > > > > > Based on review history, I'd propose Matt Butrovich and Tim Saucer as > > two good candidates, but this would be completely up to the > `iceberg-rust` > > community. > > > > > > On 2026/08/26 17:56:19 Shawn Chang wrote: > > > > Hi all, > > > > > > > > Summarizing the discussion so far, including a few points raised in > > today’s > > > > community sync. > > > > > > > > There seems to be general agreement that the current DataFusion > > integration > > > > has a velocity/reviewer bandwidth problem, and moving it to a > separate > > > > repository could help DataFusion contributors iterate more > > independently. > > > > At the same time, several open questions remain: > > > > > > > > - > > > > > > > > Whether repo separation is the right solution, versus expanding > > > > DataFusion reviewer/committer participation in iceberg-rust. > > > > - > > > > > > > > Where the boundary should be between engine specific integration > and > > > > Iceberg core functionality, and how to avoid duplicated or forked > > Iceberg > > > > implementations. (how to avoid the case where Iceberg-datafusion > > moving > > > > much faster than the core and eventually need a forked core API > > > > implementation) > > > > - > > > > > > > > How compatibility and end-to-end correctness testing should work > > across > > > > repositories, since iceberg-rust still relies on DataFusion for > > integration > > > > testing. > > > > - > > > > > > > > Where the integration should live and how to keep it under Apache > > > > governance while making it easy for both Iceberg and DataFusion > > > > contributors to maintain. > > > > > > > > So I think the main question is not only whether to move the code, > but > > how > > > > to improve development velocity without losing the close design, > > testing, > > > > and governance relationship between the engine integration and > > iceberg-rust > > > > core. > > > > > > > > > > > > Best, > > > > > > > > Shawn > > > > > > > > On Mon, Aug 24, 2026 at 4:14 AM Manu Zhang <[email protected]> > > wrote: > > > > > > > > > +1 moving the DataFusion integration and tests into a separate > > > > > apache-governed repository. Can we bring this discussion to the > > Community > > > > > Sync[1] this week? > > > > > > > > > > 1. > > > > > > > > https://docs.google.com/document/d/1YuGhUdukLP5gGiqCbk0A5_Wifqe2CZWgOd3TbhY3UQg/edit?tab=t.0 > > > > > > > > > > On Mon, Aug 24, 2026 at 3:37 PM Renjie Liu < > [email protected]> > > > > > wrote: > > > > > > > > > >> Hi, Matt: > > > > >> > > > > >> Thanks for raising this. > > > > >> > > > > >> 1) Apache governance: > > > > >> > > > > >> I would +1 for putting this in an apache repo, for example a sub > > repo of > > > > >> datafusion project. Shawn has stated most of the reasons, so I > > don't want > > > > >> to repeat it again. > > > > >> > > > > >> 2) Where do tests live/how tests should be maintained? > > > > >> > > > > >> Initially I was thinking about putting sqllogictest in > > iceberg-rust, but > > > > >> after second thought I'm leaning towards to put it in the new repo > > for two > > > > >> reasons: > > > > >> 1. It would be easier for developer of the datafusion-iceberg > > integration > > > > >> to add tests > > > > >> 2. It would make the dependency graph and version release easier. > > Though > > > > >> the dependency is on crate level rather than repo level, the > > bi-direction > > > > >> dependency may make version management weird and difficult. > > > > >> > > > > >> The downside of this approach is that it's a little unfriendly for > > > > >> iceberg-rust developers, but I think it's less frequent for > > iceberg-rust > > > > >> developers to add sqllogictests compared with datafusion-iceberg > > developers. > > > > >> > > > > >> 3) Dropping DataFusion dependencies from pyiceberg-core binding > > > > >> > > > > >> I'm not quite familiar with this part, but it sounds reasonable to > > me. > > > > >> > > > > >> > > > > >> On Sat, Aug 22, 2026 at 6:50 AM Shawn Chang < > [email protected] > > > > > > > >> wrote: > > > > >> > > > > >>> Hi Matt, > > > > >>> > > > > >>> Thanks for raising this! I generally agree that we can move the > > > > >>> datafusion > > > > >>> integration to a separate repo if that helps more DataFusion > > experts to > > > > >>> work on the integrations > > > > >>> > > > > >>> On the three considerations: > > > > >>> 1) Apache governance: > > > > >>> I think this is my biggest concern so far. This could not only > > affect > > > > >>> contributors, but also whether the downstream users could > continue > > to use > > > > >>> the integration. > > > > >>> Even if the code remains Apache-licensed, governance, release of > > > > >>> artifacts, > > > > >>> and contribution policies can matter for adoption. > > > > >>> We should explore more about the option to keep it under an > > > > >>> Apache-governed > > > > >>> repository. > > > > >>> > > > > >>> 2) Where do tests live/how tests should be maintained? > > > > >>> > > > > >>> I think this is closely related to the governance question. To > me, > > the > > > > >>> broader question is: *how do we make it easy for people from both > > the > > > > >>> Iceberg Rust and DataFusion communities to maintain the > > integration?* > > > > >>> > > > > >>> If we move it to a non-Apache repository, I worry that it could > > > > >>> eventually > > > > >>> look somewhat like the current Iceberg Java <> Trino integration: > > the > > > > >>> integration primarily lives on the Trino side and is therefore > > mostly > > > > >>> maintained by people who are already deeply involved in Trino. > > > > >>> > > > > >>> There is an important difference here, though. For Iceberg Java, > > Spark is > > > > >>> arguably the primary engine integration and has a large Iceberg > > > > >>> contributor > > > > >>> base around it, so having the Trino integration maintained more > > > > >>> independently is relatively natural. In Iceberg Rust today, > > DataFusion > > > > >>> has > > > > >>> a much more central role. It is by far the most mature engine > > integration > > > > >>> in the project, is used by our SQLLogicTest infrastructure, and > > many > > > > >>> users > > > > >>> building on Iceberg Rust are also building on DataFusion. The > > overlap > > > > >>> between the two communities is therefore much larger. > > > > >>> > > > > >>> Because of that, I would prefer that extracting the integration > > does not > > > > >>> turn it into something that is effectively owned only by the > > DataFusion > > > > >>> side. Ideally, both Iceberg Rust contributors and DataFusion > > contributors > > > > >>> should be able to review changes, maintain compatibility, and > > evolve the > > > > >>> integration together. > > > > >>> > > > > >>> 3) Dropping DataFusion dependencies from pyiceberg-core binding: > I > > think > > > > >>> it > > > > >>> makes sense to make things simpler and have left a comment on the > > github > > > > >>> issue with more detailed thoughts: > > > > >>> https://github.com/apache/iceberg-rust/issues/3036 > > > > >>> > > > > >>> Best, > > > > >>> Shawn > > > > >>> > > > > >>> On Fri, Aug 21, 2026 at 11:23 AM Matt Butrovich < > > [email protected]> > > > > >>> wrote: > > > > >>> > > > > >>> > Hello all, > > > > >>> > > > > > >>> > > > > > >>> > I've never tried emailing two different project lists at once, > > but it > > > > >>> was > > > > >>> > suggested that I do so to try to track the conversation across > > both > > > > >>> > communities. We'll see how this threads on the mailing lists. > > > > >>> > > > > > >>> > > > > > >>> > There has been a GitHub Discussion > > > > >>> > > > > > >>> > > > https://github.com/apache/iceberg-rust/discussions/2992#discussioncomment-18082533 > > > > >>> , > > > > >>> > a GitHub Issue > > https://github.com/apache/iceberg-rust/issues/3029, and > > > > >>> > it's been a long topic of conversation in the past two weeks in > > both > > > > >>> the > > > > >>> > DataFusion and Iceberg Rust Community Calls to discuss moving > the > > > > >>> > DataFusion integration from Iceberg Rust to a separate > > repository. I > > > > >>> will > > > > >>> > try to summarize some of the major points as I understand them, > > but the > > > > >>> > conversations are the ground truth and please feel free to > > correct me > > > > >>> here. > > > > >>> > > > > > >>> > > > > > >>> > The DataFusion TableProvider integration in Iceberg Rust makes > > > > >>> DataFusion > > > > >>> > a dependency for Iceberg Rust. The integration exists for > > multiple > > > > >>> reasons: > > > > >>> > > > > > >>> > 1) an engine to execute Iceberg Rust's corpus of sqllogictest > > files > > > > >>> > > > > > >>> > 2) a TableProvider integration for DataFusion users to interact > > with > > > > >>> > Iceberg tables > > > > >>> > > > > > >>> > > > > > >>> > Some of the motivations to break out the integration: > > > > >>> > > > > > >>> > 1) There have been a number of issues and pull requests against > > the > > > > >>> > DataFusion TableProvider in Iceberg Rust as users want to add > > more > > > > >>> > features, and they often go stale. I don't believe there are > many > > > > >>> > committers/PMC members familiar with or using the DataFusion > > > > >>> integration. > > > > >>> > > > > > >>> > 2) Iceberg Rust would like to stay as engine-agnostic as > > possible. A > > > > >>> > recent DataFusion Ballista integration was declined for this > > reason > > > > >>> > https://github.com/apache/iceberg-rust/pull/2613. > > > > >>> > > > > > >>> > 3) Other projects that rely on both Iceberg Rust and DataFusion > > (e.g., > > > > >>> > Comet) are blocked by Iceberg Rust upgrading its DataFusion and > > Arrow > > > > >>> > dependencies before they can upgrade. > > > > >>> > > > > > >>> > > > > > >>> > Some considerations for both communities: > > > > >>> > > > > > >>> > 1) Where would this DataFusion TableProvider live? Most > > specifically, > > > > >>> > would it be an Apache-governed project? As folks like > @andygrove > > point > > > > >>> out, > > > > >>> > this can affect whether some community members could contribute > > to it. > > > > >>> > There is a datafusion-contrib org for DataFusion-related > > projects to > > > > >>> have > > > > >>> > visibility but no Apache governance, but there may be options > to > > put it > > > > >>> > under an Apache repository. > > > > >>> > > > > > >>> > 2) How would Iceberg Rust continue to run sqllogictests for > > regression > > > > >>> > testing? Does this live in a different repository that depends > > on this > > > > >>> new > > > > >>> > Iceberg Rust TableProvider crate? Would we be able to test Pull > > > > >>> Requests on > > > > >>> > Iceberg Rust with sqllogictests? > > > > >>> > > > > > >>> > 3) @kevinqliu is familiar with the Python bindings in Iceberg > > Rust, > > > > >>> and I > > > > >>> > believe that DataFusion dependency might be removed as well, > but > > that > > > > >>> is > > > > >>> > not the core focus of this conversation. He has also proposed > > removing > > > > >>> that > > > > >>> > and using the DataFusion Python bindings directly. > > > > >>> > > > > > >>> > > > > > >>> > I'm sure I'm forgetting things, but this email is long enough. > > > > >>> > > > > > >>> > > > > > >>> > Thanks everyone for the discussion thus far, and looking > forward > > to > > > > >>> more > > > > >>> > input on this. > > > > >>> > > > > > >>> > > > > > >>> > -Matt > > > > >>> > > > > > >>> > > > > > >>> > > > > > >>> > > > > > >>> > > > > >> > > > > > > > > > > Xuanwo > > > > > > https://xuanwo.io/ > > > > > >
