+ dev@iceberg (lesson learned, emailing 2 devlist might not be the best idea logistically haha)
On Fri, Sep 4, 2026 at 10:48 AM Kevin Liu <[email protected]> wrote: > Thanks for all the great discussions so far. I'm glad we found a solution > that benefits the broader ecosystem. > I'm +1 to moving to an apache-governed location, > apache/datafusion-iceberg seems like a great place. I like the process > Andrew proposed, happy to help with the logistics. > > Best, > Kevin Liu > > PS I have also removed the iceberg python binding that exports > DataFusion's TableProvider, so that's 1 less thing we have to worry about. > https://github.com/apache/iceberg-rust/issues/3036 > > On Fri, Sep 4, 2026 at 6:28 AM Gabriel Musat <[email protected]> wrote: > >> Sounds like a good direction, +1 (non binding) >> >> I can help porting the commit history of the DataFusion-Iceberg >> integration to the new repo if PMCs agree. >> >> On 2026/09/04 12:27:30 Andrew Lamb wrote: >> > I think this is a great idea as well -- thank you for bringing it up and >> > for the great discussions so far. >> > >> > While at VLDB this past week, I spoke to at least three people from >> > companies adding Apache Iceberg support to their products. All of them >> had >> > to fork iceberg-rust for one reason or another. >> > >> > I think we have a huge need to improve our ability to work together and >> > accelerate everyone's efforts. Giving the DataFusion integration access >> to >> > more expert maintainers with bandwidth I think will help everyone. >> > >> > There seems to be consensus that we should move the integration code to >> > DataFusion governance, and that it should remain in the ASF, though some >> > open technical questions remain. >> > >> > If that is the case, I propose the following specific process: >> > 1. Create the new gitub repository in Apache for the code (e.g. >> > apache/datafusion-iceberg) >> > 2. Create a PR in the new repo with the proposed code >> > 3. Hold a formal vote on the iceberg dev list to move the integration >> code >> > to DataFusion >> > 4. Hold a formal vote on the DataFusion dev list to accept the new code >> > >> > I am happy to help with the logistics (e.g. ASF INFRA ticket to create >> the >> > new repo, votes, etc) but I am not expert enough to create the proposed >> PR. >> > >> > Please let me know your thoughts, >> > Andrew >> > (PMC Chair of DataFusion) >> > >> > On Tue, Sep 1, 2026 at 9:48 PM Renjie Liu <[email protected]> >> wrote: >> > >> > > To add more background about the relationship between comet and >> datafusion >> > > for those who are not familiar with them. >> > > >> > > Apache datafusion is a popular extensible compute engine written in >> rust. >> > > Apache comet is an apache spark accelerator builton on apache >> datafusion, >> > > and also a subproject of apache datafusion. >> > > >> > > datafusion-iceberg is an apache datafusion extension built on >> iceberg-rust, >> > > and comet's iceberg support is built on it. >> > > >> > > On Tue, Sep 1, 2026 at 9:39 PM Andy Grove <[email protected]> >> wrote: >> > > >> > > > +1 for moving to DataFusion PMC. The iceberg-datafusion integration >> is >> > > > very important for Comet. >> > > > >> > > > On Mon, Aug 31, 2026 at 11:42 PM Xuanwo <[email protected]> wrote: >> > > > > >> > > > > TBH, I also support moving to the DataFusion PMC. >> > > > > >> > > > > - DataFusion is the largest dependency in iceberg-datafusion. >> > > > > - The largest downstream user of iceberg-datafusion is Comet, >> which >> > > > shares many of the same PMC members as DataFusion. >> > > > > >> > > > > It feels natural to be part of the DataFusion PMC. As long as DF >> PMC is >> > > > willing to accept this project, it LGTM. >> > > > > >> > > > > On Tue, Sep 1, 2026, at 12:03, Renjie Liu wrote: >> > > > > >> > > > > I don't think the testing should be a blocker of moving >> > > > iceberg-datafusion out of iceberg-rust repo. From what I learn, >> most of >> > > the >> > > > sqllogictests are in pr of modifying iceberg-datafusion integration, >> > > there >> > > > are only few cases where we rely on sqllogictests to verify >> features. >> > > > > >> > > > > > How would that sound for the `iceberg-rust` community to grant a >> > > > couple of Apache DataFusion PMCs committer access to the repository, >> > > > limited to the `iceberg-datafusion` crate via a CODEOWNERS file, so >> that >> > > > they can help push reviews and PRs forward independently? >> > > > > >> > > > > I'm not sure if this is feasible, but moving the >> iceberg-datafusion >> > > > crate to apache datafusion project sounds a more reasonable >> approach to >> > > me. >> > > > It's still governed by Apache, and most of the code is related to >> > > > DataFusion, so I think the DataFusion community is in a better >> position >> > > to >> > > > define the vision and design for it. >> > > > > >> > > > > On Mon, Aug 31, 2026 at 10:45 PM Gabriel Musat < >> [email protected]> >> > > > wrote: >> > > > > >> > > > > Hi everyone, >> > > > > >> > > > > Based on these two facts: >> > > > > - There's a current reviewer bandwidth problem that hurts >> development >> > > > velocity of the `iceberg-datafusion` crate currently hosted under >> > > > `apache/iceberg-rust`. >> > > > > - There is value in maintaining the `iceberg-datafusion` crate >> inside >> > > > `apache/iceberg-rust` for testing, design and governance. >> > > > > >> > > > > How would that sound for the `iceberg-rust` community to grant a >> couple >> > > > of Apache DataFusion PMCs committer access to the repository, >> limited to >> > > > the `iceberg-datafusion` crate via a CODEOWNERS file, so that they >> can >> > > help >> > > > push reviews and PRs forward independently? >> > > > > >> > > > > Based on review history, I'd propose Matt Butrovich and Tim >> Saucer as >> > > > two good candidates, but this would be completely up to the >> > > `iceberg-rust` >> > > > community. >> > > > > >> > > > > On 2026/08/26 17:56:19 Shawn Chang wrote: >> > > > > > Hi all, >> > > > > > >> > > > > > Summarizing the discussion so far, including a few points >> raised in >> > > > today’s >> > > > > > community sync. >> > > > > > >> > > > > > There seems to be general agreement that the current DataFusion >> > > > integration >> > > > > > has a velocity/reviewer bandwidth problem, and moving it to a >> > > separate >> > > > > > repository could help DataFusion contributors iterate more >> > > > independently. >> > > > > > At the same time, several open questions remain: >> > > > > > >> > > > > > - >> > > > > > >> > > > > > Whether repo separation is the right solution, versus >> expanding >> > > > > > DataFusion reviewer/committer participation in iceberg-rust. >> > > > > > - >> > > > > > >> > > > > > Where the boundary should be between engine specific >> integration >> > > and >> > > > > > Iceberg core functionality, and how to avoid duplicated or >> forked >> > > > Iceberg >> > > > > > implementations. (how to avoid the case where >> Iceberg-datafusion >> > > > moving >> > > > > > much faster than the core and eventually need a forked core >> API >> > > > > > implementation) >> > > > > > - >> > > > > > >> > > > > > How compatibility and end-to-end correctness testing should >> work >> > > > across >> > > > > > repositories, since iceberg-rust still relies on DataFusion >> for >> > > > integration >> > > > > > testing. >> > > > > > - >> > > > > > >> > > > > > Where the integration should live and how to keep it under >> Apache >> > > > > > governance while making it easy for both Iceberg and >> DataFusion >> > > > > > contributors to maintain. >> > > > > > >> > > > > > So I think the main question is not only whether to move the >> code, >> > > but >> > > > how >> > > > > > to improve development velocity without losing the close design, >> > > > testing, >> > > > > > and governance relationship between the engine integration and >> > > > iceberg-rust >> > > > > > core. >> > > > > > >> > > > > > >> > > > > > Best, >> > > > > > >> > > > > > Shawn >> > > > > > >> > > > > > On Mon, Aug 24, 2026 at 4:14 AM Manu Zhang < >> [email protected]> >> > > > wrote: >> > > > > > >> > > > > > > +1 moving the DataFusion integration and tests into a separate >> > > > > > > apache-governed repository. Can we bring this discussion to >> the >> > > > Community >> > > > > > > Sync[1] this week? >> > > > > > > >> > > > > > > 1. >> > > > > > > >> > > > >> > > >> https://docs.google.com/document/d/1YuGhUdukLP5gGiqCbk0A5_Wifqe2CZWgOd3TbhY3UQg/edit?tab=t.0 >> > > > > > > >> > > > > > > On Mon, Aug 24, 2026 at 3:37 PM Renjie Liu < >> > > [email protected]> >> > > > > > > wrote: >> > > > > > > >> > > > > > >> Hi, Matt: >> > > > > > >> >> > > > > > >> Thanks for raising this. >> > > > > > >> >> > > > > > >> 1) Apache governance: >> > > > > > >> >> > > > > > >> I would +1 for putting this in an apache repo, for example a >> sub >> > > > repo of >> > > > > > >> datafusion project. Shawn has stated most of the reasons, so >> I >> > > > don't want >> > > > > > >> to repeat it again. >> > > > > > >> >> > > > > > >> 2) Where do tests live/how tests should be maintained? >> > > > > > >> >> > > > > > >> Initially I was thinking about putting sqllogictest in >> > > > iceberg-rust, but >> > > > > > >> after second thought I'm leaning towards to put it in the >> new repo >> > > > for two >> > > > > > >> reasons: >> > > > > > >> 1. It would be easier for developer of the datafusion-iceberg >> > > > integration >> > > > > > >> to add tests >> > > > > > >> 2. It would make the dependency graph and version release >> easier. >> > > > Though >> > > > > > >> the dependency is on crate level rather than repo level, the >> > > > bi-direction >> > > > > > >> dependency may make version management weird and difficult. >> > > > > > >> >> > > > > > >> The downside of this approach is that it's a little >> unfriendly for >> > > > > > >> iceberg-rust developers, but I think it's less frequent for >> > > > iceberg-rust >> > > > > > >> developers to add sqllogictests compared with >> datafusion-iceberg >> > > > developers. >> > > > > > >> >> > > > > > >> 3) Dropping DataFusion dependencies from pyiceberg-core >> binding >> > > > > > >> >> > > > > > >> I'm not quite familiar with this part, but it sounds >> reasonable to >> > > > me. >> > > > > > >> >> > > > > > >> >> > > > > > >> On Sat, Aug 22, 2026 at 6:50 AM Shawn Chang < >> > > [email protected] >> > > > > >> > > > > > >> wrote: >> > > > > > >> >> > > > > > >>> Hi Matt, >> > > > > > >>> >> > > > > > >>> Thanks for raising this! I generally agree that we can move >> the >> > > > > > >>> datafusion >> > > > > > >>> integration to a separate repo if that helps more DataFusion >> > > > experts to >> > > > > > >>> work on the integrations >> > > > > > >>> >> > > > > > >>> On the three considerations: >> > > > > > >>> 1) Apache governance: >> > > > > > >>> I think this is my biggest concern so far. This could not >> only >> > > > affect >> > > > > > >>> contributors, but also whether the downstream users could >> > > continue >> > > > to use >> > > > > > >>> the integration. >> > > > > > >>> Even if the code remains Apache-licensed, governance, >> release of >> > > > > > >>> artifacts, >> > > > > > >>> and contribution policies can matter for adoption. >> > > > > > >>> We should explore more about the option to keep it under an >> > > > > > >>> Apache-governed >> > > > > > >>> repository. >> > > > > > >>> >> > > > > > >>> 2) Where do tests live/how tests should be maintained? >> > > > > > >>> >> > > > > > >>> I think this is closely related to the governance question. >> To >> > > me, >> > > > the >> > > > > > >>> broader question is: *how do we make it easy for people >> from both >> > > > the >> > > > > > >>> Iceberg Rust and DataFusion communities to maintain the >> > > > integration?* >> > > > > > >>> >> > > > > > >>> If we move it to a non-Apache repository, I worry that it >> could >> > > > > > >>> eventually >> > > > > > >>> look somewhat like the current Iceberg Java <> Trino >> integration: >> > > > the >> > > > > > >>> integration primarily lives on the Trino side and is >> therefore >> > > > mostly >> > > > > > >>> maintained by people who are already deeply involved in >> Trino. >> > > > > > >>> >> > > > > > >>> There is an important difference here, though. For Iceberg >> Java, >> > > > Spark is >> > > > > > >>> arguably the primary engine integration and has a large >> Iceberg >> > > > > > >>> contributor >> > > > > > >>> base around it, so having the Trino integration maintained >> more >> > > > > > >>> independently is relatively natural. In Iceberg Rust today, >> > > > DataFusion >> > > > > > >>> has >> > > > > > >>> a much more central role. It is by far the most mature >> engine >> > > > integration >> > > > > > >>> in the project, is used by our SQLLogicTest infrastructure, >> and >> > > > many >> > > > > > >>> users >> > > > > > >>> building on Iceberg Rust are also building on DataFusion. >> The >> > > > overlap >> > > > > > >>> between the two communities is therefore much larger. >> > > > > > >>> >> > > > > > >>> Because of that, I would prefer that extracting the >> integration >> > > > does not >> > > > > > >>> turn it into something that is effectively owned only by the >> > > > DataFusion >> > > > > > >>> side. Ideally, both Iceberg Rust contributors and DataFusion >> > > > contributors >> > > > > > >>> should be able to review changes, maintain compatibility, >> and >> > > > evolve the >> > > > > > >>> integration together. >> > > > > > >>> >> > > > > > >>> 3) Dropping DataFusion dependencies from pyiceberg-core >> binding: >> > > I >> > > > think >> > > > > > >>> it >> > > > > > >>> makes sense to make things simpler and have left a comment >> on the >> > > > github >> > > > > > >>> issue with more detailed thoughts: >> > > > > > >>> https://github.com/apache/iceberg-rust/issues/3036 >> > > > > > >>> >> > > > > > >>> Best, >> > > > > > >>> Shawn >> > > > > > >>> >> > > > > > >>> On Fri, Aug 21, 2026 at 11:23 AM Matt Butrovich < >> > > > [email protected]> >> > > > > > >>> wrote: >> > > > > > >>> >> > > > > > >>> > Hello all, >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > I've never tried emailing two different project lists at >> once, >> > > > but it >> > > > > > >>> was >> > > > > > >>> > suggested that I do so to try to track the conversation >> across >> > > > both >> > > > > > >>> > communities. We'll see how this threads on the mailing >> lists. >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > There has been a GitHub Discussion >> > > > > > >>> > >> > > > > > >>> >> > > > >> > > >> https://github.com/apache/iceberg-rust/discussions/2992#discussioncomment-18082533 >> > > > > > >>> , >> > > > > > >>> > a GitHub Issue >> > > > https://github.com/apache/iceberg-rust/issues/3029, and >> > > > > > >>> > it's been a long topic of conversation in the past two >> weeks in >> > > > both >> > > > > > >>> the >> > > > > > >>> > DataFusion and Iceberg Rust Community Calls to discuss >> moving >> > > the >> > > > > > >>> > DataFusion integration from Iceberg Rust to a separate >> > > > repository. I >> > > > > > >>> will >> > > > > > >>> > try to summarize some of the major points as I understand >> them, >> > > > but the >> > > > > > >>> > conversations are the ground truth and please feel free to >> > > > correct me >> > > > > > >>> here. >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > The DataFusion TableProvider integration in Iceberg Rust >> makes >> > > > > > >>> DataFusion >> > > > > > >>> > a dependency for Iceberg Rust. The integration exists for >> > > > multiple >> > > > > > >>> reasons: >> > > > > > >>> > >> > > > > > >>> > 1) an engine to execute Iceberg Rust's corpus of >> sqllogictest >> > > > files >> > > > > > >>> > >> > > > > > >>> > 2) a TableProvider integration for DataFusion users to >> interact >> > > > with >> > > > > > >>> > Iceberg tables >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > Some of the motivations to break out the integration: >> > > > > > >>> > >> > > > > > >>> > 1) There have been a number of issues and pull requests >> against >> > > > the >> > > > > > >>> > DataFusion TableProvider in Iceberg Rust as users want to >> add >> > > > more >> > > > > > >>> > features, and they often go stale. I don't believe there >> are >> > > many >> > > > > > >>> > committers/PMC members familiar with or using the >> DataFusion >> > > > > > >>> integration. >> > > > > > >>> > >> > > > > > >>> > 2) Iceberg Rust would like to stay as engine-agnostic as >> > > > possible. A >> > > > > > >>> > recent DataFusion Ballista integration was declined for >> this >> > > > reason >> > > > > > >>> > https://github.com/apache/iceberg-rust/pull/2613. >> > > > > > >>> > >> > > > > > >>> > 3) Other projects that rely on both Iceberg Rust and >> DataFusion >> > > > (e.g., >> > > > > > >>> > Comet) are blocked by Iceberg Rust upgrading its >> DataFusion and >> > > > Arrow >> > > > > > >>> > dependencies before they can upgrade. >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > Some considerations for both communities: >> > > > > > >>> > >> > > > > > >>> > 1) Where would this DataFusion TableProvider live? Most >> > > > specifically, >> > > > > > >>> > would it be an Apache-governed project? As folks like >> > > @andygrove >> > > > point >> > > > > > >>> out, >> > > > > > >>> > this can affect whether some community members could >> contribute >> > > > to it. >> > > > > > >>> > There is a datafusion-contrib org for DataFusion-related >> > > > projects to >> > > > > > >>> have >> > > > > > >>> > visibility but no Apache governance, but there may be >> options >> > > to >> > > > put it >> > > > > > >>> > under an Apache repository. >> > > > > > >>> > >> > > > > > >>> > 2) How would Iceberg Rust continue to run sqllogictests >> for >> > > > regression >> > > > > > >>> > testing? Does this live in a different repository that >> depends >> > > > on this >> > > > > > >>> new >> > > > > > >>> > Iceberg Rust TableProvider crate? Would we be able to >> test Pull >> > > > > > >>> Requests on >> > > > > > >>> > Iceberg Rust with sqllogictests? >> > > > > > >>> > >> > > > > > >>> > 3) @kevinqliu is familiar with the Python bindings in >> Iceberg >> > > > Rust, >> > > > > > >>> and I >> > > > > > >>> > believe that DataFusion dependency might be removed as >> well, >> > > but >> > > > that >> > > > > > >>> is >> > > > > > >>> > not the core focus of this conversation. He has also >> proposed >> > > > removing >> > > > > > >>> that >> > > > > > >>> > and using the DataFusion Python bindings directly. >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > I'm sure I'm forgetting things, but this email is long >> enough. >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > Thanks everyone for the discussion thus far, and looking >> > > forward >> > > > to >> > > > > > >>> more >> > > > > > >>> > input on this. >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > -Matt >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> > >> > > > > > >>> >> > > > > > >> >> > > > > > >> > > > > >> > > > > Xuanwo >> > > > > >> > > > > https://xuanwo.io/ >> > > > > >> > > > >> > > >> > >> >> --------------------------------------------------------------------- >> To unsubscribe, e-mail: [email protected] >> For additional commands, e-mail: [email protected] >> >>
