Sounds like a good direction, +1 (non binding)

I can help porting the commit history of the DataFusion-Iceberg integration to 
the new repo if PMCs agree.

On 2026/09/04 12:27:30 Andrew Lamb wrote:
> I think this is a great idea as well -- thank you for bringing it up and
> for the great discussions so far.
> 
> While at VLDB this past week, I spoke to at least three people from
> companies adding Apache Iceberg support to their products. All of them had
> to fork iceberg-rust for one reason or another.
> 
> I think we have a huge need to improve our ability to work together and
> accelerate everyone's efforts. Giving the DataFusion integration access to
> more expert maintainers with bandwidth I think will help everyone.
> 
> There seems to be consensus that we should move the integration code to
> DataFusion governance, and that it should remain in the ASF, though some
> open technical questions remain.
> 
> If that is the case, I propose the following specific process:
> 1. Create the new gitub repository in Apache for the code (e.g.
> apache/datafusion-iceberg)
> 2. Create a PR in the new repo with the proposed code
> 3. Hold a formal vote on the iceberg dev list to move the integration code
> to DataFusion
> 4. Hold a formal vote on the DataFusion dev list to accept the new code
> 
> I am happy to help with the logistics (e.g. ASF INFRA ticket to create the
> new repo, votes, etc) but I am not expert enough to create the proposed PR.
> 
> Please let me know your thoughts,
> Andrew
> (PMC Chair of DataFusion)
> 
> On Tue, Sep 1, 2026 at 9:48 PM Renjie Liu <[email protected]> wrote:
> 
> > To add more background about the relationship between comet and datafusion
> > for those who are not familiar with them.
> >
> > Apache datafusion is a popular extensible compute engine written in rust.
> > Apache comet is an apache spark accelerator builton on apache datafusion,
> > and also a subproject of apache datafusion.
> >
> > datafusion-iceberg is an apache datafusion extension built on iceberg-rust,
> > and comet's iceberg support is built on it.
> >
> > On Tue, Sep 1, 2026 at 9:39 PM Andy Grove <[email protected]> wrote:
> >
> > > +1 for moving to DataFusion PMC. The iceberg-datafusion integration is
> > > very important for Comet.
> > >
> > > On Mon, Aug 31, 2026 at 11:42 PM Xuanwo <[email protected]> wrote:
> > > >
> > > > TBH, I also support moving to the DataFusion PMC.
> > > >
> > > > - DataFusion is the largest dependency in iceberg-datafusion.
> > > > - The largest downstream user of iceberg-datafusion is Comet, which
> > > shares many of the same PMC members as DataFusion.
> > > >
> > > > It feels natural to be part of the DataFusion PMC. As long as DF PMC is
> > > willing to accept this project, it LGTM.
> > > >
> > > > On Tue, Sep 1, 2026, at 12:03, Renjie Liu wrote:
> > > >
> > > > I don't think the testing should be a blocker of moving
> > > iceberg-datafusion out of iceberg-rust repo. From what I learn, most of
> > the
> > > sqllogictests are in pr of modifying iceberg-datafusion integration,
> > there
> > > are only few cases where we rely on sqllogictests to verify features.
> > > >
> > > > > How would that sound for the `iceberg-rust` community to grant a
> > > couple of Apache DataFusion PMCs committer access to the repository,
> > > limited to the `iceberg-datafusion` crate via a CODEOWNERS file, so that
> > > they can help push reviews and PRs forward independently?
> > > >
> > > > I'm not sure if this is feasible, but moving the iceberg-datafusion
> > > crate to apache datafusion project sounds a more reasonable approach to
> > me.
> > > It's still governed by Apache, and most of the code is related to
> > > DataFusion, so I think the DataFusion community is in a better position
> > to
> > > define the vision and design for it.
> > > >
> > > > On Mon, Aug 31, 2026 at 10:45 PM Gabriel Musat <[email protected]>
> > > wrote:
> > > >
> > > > Hi everyone,
> > > >
> > > > Based on these two facts:
> > > > - There's a current reviewer bandwidth problem that hurts development
> > > velocity of the `iceberg-datafusion` crate currently hosted under
> > > `apache/iceberg-rust`.
> > > > - There is value in maintaining the `iceberg-datafusion` crate inside
> > > `apache/iceberg-rust` for testing, design and governance.
> > > >
> > > > How would that sound for the `iceberg-rust` community to grant a couple
> > > of Apache DataFusion PMCs committer access to the repository, limited to
> > > the `iceberg-datafusion` crate via a CODEOWNERS file, so that they can
> > help
> > > push reviews and PRs forward independently?
> > > >
> > > > Based on review history, I'd propose Matt Butrovich and Tim Saucer as
> > > two good candidates, but this would be completely up to the
> > `iceberg-rust`
> > > community.
> > > >
> > > > On 2026/08/26 17:56:19 Shawn Chang wrote:
> > > > > Hi all,
> > > > >
> > > > > Summarizing the discussion so far, including a few points raised in
> > > today’s
> > > > > community sync.
> > > > >
> > > > > There seems to be general agreement that the current DataFusion
> > > integration
> > > > > has a velocity/reviewer bandwidth problem, and moving it to a
> > separate
> > > > > repository could help DataFusion contributors iterate more
> > > independently.
> > > > > At the same time, several open questions remain:
> > > > >
> > > > >    -
> > > > >
> > > > >    Whether repo separation is the right solution, versus expanding
> > > > >    DataFusion reviewer/committer participation in iceberg-rust.
> > > > >    -
> > > > >
> > > > >    Where the boundary should be between engine specific integration
> > and
> > > > >    Iceberg core functionality, and how to avoid duplicated or forked
> > > Iceberg
> > > > >    implementations. (how to avoid the case where Iceberg-datafusion
> > > moving
> > > > >    much faster than the core and eventually need a forked core API
> > > > >    implementation)
> > > > >    -
> > > > >
> > > > >    How compatibility and end-to-end correctness testing should work
> > > across
> > > > >    repositories, since iceberg-rust still relies on DataFusion for
> > > integration
> > > > >    testing.
> > > > >    -
> > > > >
> > > > >    Where the integration should live and how to keep it under Apache
> > > > >    governance while making it easy for both Iceberg and DataFusion
> > > > >    contributors to maintain.
> > > > >
> > > > > So I think the main question is not only whether to move the code,
> > but
> > > how
> > > > > to improve development velocity without losing the close design,
> > > testing,
> > > > > and governance relationship between the engine integration and
> > > iceberg-rust
> > > > > core.
> > > > >
> > > > >
> > > > > Best,
> > > > >
> > > > > Shawn
> > > > >
> > > > > On Mon, Aug 24, 2026 at 4:14 AM Manu Zhang <[email protected]>
> > > wrote:
> > > > >
> > > > > > +1 moving the DataFusion integration and tests into a separate
> > > > > > apache-governed repository. Can we bring this discussion to the
> > > Community
> > > > > > Sync[1] this week?
> > > > > >
> > > > > > 1.
> > > > > >
> > >
> > https://docs.google.com/document/d/1YuGhUdukLP5gGiqCbk0A5_Wifqe2CZWgOd3TbhY3UQg/edit?tab=t.0
> > > > > >
> > > > > > On Mon, Aug 24, 2026 at 3:37 PM Renjie Liu <
> > [email protected]>
> > > > > > wrote:
> > > > > >
> > > > > >> Hi, Matt:
> > > > > >>
> > > > > >> Thanks for raising this.
> > > > > >>
> > > > > >> 1) Apache governance:
> > > > > >>
> > > > > >> I would +1 for putting this in an apache repo, for example a sub
> > > repo of
> > > > > >> datafusion project. Shawn has stated most of the reasons, so I
> > > don't want
> > > > > >> to repeat it again.
> > > > > >>
> > > > > >> 2) Where do tests live/how tests should be maintained?
> > > > > >>
> > > > > >> Initially I was thinking about putting sqllogictest in
> > > iceberg-rust, but
> > > > > >> after second thought I'm leaning towards to put it in the new repo
> > > for two
> > > > > >> reasons:
> > > > > >> 1. It would be easier for developer of the datafusion-iceberg
> > > integration
> > > > > >> to add tests
> > > > > >> 2. It would make the dependency graph and version release easier.
> > > Though
> > > > > >> the dependency is on crate level rather than repo level, the
> > > bi-direction
> > > > > >> dependency may make version management weird and difficult.
> > > > > >>
> > > > > >> The downside of this approach is that it's a little unfriendly for
> > > > > >> iceberg-rust developers, but I think it's less frequent for
> > > iceberg-rust
> > > > > >> developers to add sqllogictests compared with datafusion-iceberg
> > > developers.
> > > > > >>
> > > > > >> 3) Dropping DataFusion dependencies from pyiceberg-core binding
> > > > > >>
> > > > > >> I'm not quite familiar with this part, but it sounds reasonable to
> > > me.
> > > > > >>
> > > > > >>
> > > > > >> On Sat, Aug 22, 2026 at 6:50 AM Shawn Chang <
> > [email protected]
> > > >
> > > > > >> wrote:
> > > > > >>
> > > > > >>> Hi Matt,
> > > > > >>>
> > > > > >>> Thanks for raising this! I generally agree that we can move the
> > > > > >>> datafusion
> > > > > >>> integration to a separate repo if that helps more DataFusion
> > > experts to
> > > > > >>> work on the integrations
> > > > > >>>
> > > > > >>> On the three considerations:
> > > > > >>> 1) Apache governance:
> > > > > >>> I think this is my biggest concern so far. This could not only
> > > affect
> > > > > >>> contributors, but also whether the downstream users could
> > continue
> > > to use
> > > > > >>> the integration.
> > > > > >>> Even if the code remains Apache-licensed, governance, release of
> > > > > >>> artifacts,
> > > > > >>> and contribution policies can matter for adoption.
> > > > > >>> We should explore more about the option to keep it under an
> > > > > >>> Apache-governed
> > > > > >>> repository.
> > > > > >>>
> > > > > >>> 2) Where do tests live/how tests should be maintained?
> > > > > >>>
> > > > > >>> I think this is closely related to the governance question. To
> > me,
> > > the
> > > > > >>> broader question is: *how do we make it easy for people from both
> > > the
> > > > > >>> Iceberg Rust and DataFusion communities to maintain the
> > > integration?*
> > > > > >>>
> > > > > >>> If we move it to a non-Apache repository, I worry that it could
> > > > > >>> eventually
> > > > > >>> look somewhat like the current Iceberg Java <> Trino integration:
> > > the
> > > > > >>> integration primarily lives on the Trino side and is therefore
> > > mostly
> > > > > >>> maintained by people who are already deeply involved in Trino.
> > > > > >>>
> > > > > >>> There is an important difference here, though. For Iceberg Java,
> > > Spark is
> > > > > >>> arguably the primary engine integration and has a large Iceberg
> > > > > >>> contributor
> > > > > >>> base around it, so having the Trino integration maintained more
> > > > > >>> independently is relatively natural. In Iceberg Rust today,
> > > DataFusion
> > > > > >>> has
> > > > > >>> a much more central role. It is by far the most mature engine
> > > integration
> > > > > >>> in the project, is used by our SQLLogicTest infrastructure, and
> > > many
> > > > > >>> users
> > > > > >>> building on Iceberg Rust are also building on DataFusion. The
> > > overlap
> > > > > >>> between the two communities is therefore much larger.
> > > > > >>>
> > > > > >>> Because of that, I would prefer that extracting the integration
> > > does not
> > > > > >>> turn it into something that is effectively owned only by the
> > > DataFusion
> > > > > >>> side. Ideally, both Iceberg Rust contributors and DataFusion
> > > contributors
> > > > > >>> should be able to review changes, maintain compatibility, and
> > > evolve the
> > > > > >>> integration together.
> > > > > >>>
> > > > > >>> 3) Dropping DataFusion dependencies from pyiceberg-core binding:
> > I
> > > think
> > > > > >>> it
> > > > > >>> makes sense to make things simpler and have left a comment on the
> > > github
> > > > > >>> issue with more detailed thoughts:
> > > > > >>> https://github.com/apache/iceberg-rust/issues/3036
> > > > > >>>
> > > > > >>> Best,
> > > > > >>> Shawn
> > > > > >>>
> > > > > >>> On Fri, Aug 21, 2026 at 11:23 AM Matt Butrovich <
> > > [email protected]>
> > > > > >>> wrote:
> > > > > >>>
> > > > > >>> > Hello all,
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > I've never tried emailing two different project lists at once,
> > > but it
> > > > > >>> was
> > > > > >>> > suggested that I do so to try to track the conversation across
> > > both
> > > > > >>> > communities. We'll see how this threads on the mailing lists.
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > There has been a GitHub Discussion
> > > > > >>> >
> > > > > >>>
> > >
> > https://github.com/apache/iceberg-rust/discussions/2992#discussioncomment-18082533
> > > > > >>> ,
> > > > > >>> > a GitHub Issue
> > > https://github.com/apache/iceberg-rust/issues/3029, and
> > > > > >>> > it's been a long topic of conversation in the past two weeks in
> > > both
> > > > > >>> the
> > > > > >>> > DataFusion and Iceberg Rust Community Calls to discuss moving
> > the
> > > > > >>> > DataFusion integration from Iceberg Rust to a separate
> > > repository. I
> > > > > >>> will
> > > > > >>> > try to summarize some of the major points as I understand them,
> > > but the
> > > > > >>> > conversations are the ground truth and please feel free to
> > > correct me
> > > > > >>> here.
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > The DataFusion TableProvider integration in Iceberg Rust makes
> > > > > >>> DataFusion
> > > > > >>> > a dependency for Iceberg Rust. The integration exists for
> > > multiple
> > > > > >>> reasons:
> > > > > >>> >
> > > > > >>> > 1) an engine to execute Iceberg Rust's corpus of sqllogictest
> > > files
> > > > > >>> >
> > > > > >>> > 2) a TableProvider integration for DataFusion users to interact
> > > with
> > > > > >>> > Iceberg tables
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > Some of the motivations to break out the integration:
> > > > > >>> >
> > > > > >>> > 1) There have been a number of issues and pull requests against
> > > the
> > > > > >>> > DataFusion TableProvider in Iceberg Rust as users want to add
> > > more
> > > > > >>> > features, and they often go stale. I don't believe there are
> > many
> > > > > >>> > committers/PMC members familiar with or using the DataFusion
> > > > > >>> integration.
> > > > > >>> >
> > > > > >>> > 2) Iceberg Rust would like to stay as engine-agnostic as
> > > possible. A
> > > > > >>> > recent DataFusion Ballista integration was declined for this
> > > reason
> > > > > >>> > https://github.com/apache/iceberg-rust/pull/2613.
> > > > > >>> >
> > > > > >>> > 3) Other projects that rely on both Iceberg Rust and DataFusion
> > > (e.g.,
> > > > > >>> > Comet) are blocked by Iceberg Rust upgrading its DataFusion and
> > > Arrow
> > > > > >>> > dependencies before they can upgrade.
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > Some considerations for both communities:
> > > > > >>> >
> > > > > >>> > 1) Where would this DataFusion TableProvider live? Most
> > > specifically,
> > > > > >>> > would it be an Apache-governed project? As folks like
> > @andygrove
> > > point
> > > > > >>> out,
> > > > > >>> > this can affect whether some community members could contribute
> > > to it.
> > > > > >>> > There is a datafusion-contrib org for DataFusion-related
> > > projects to
> > > > > >>> have
> > > > > >>> > visibility but no Apache governance, but there may be options
> > to
> > > put it
> > > > > >>> > under an Apache repository.
> > > > > >>> >
> > > > > >>> > 2) How would Iceberg Rust continue to run sqllogictests for
> > > regression
> > > > > >>> > testing? Does this live in a different repository that depends
> > > on this
> > > > > >>> new
> > > > > >>> > Iceberg Rust TableProvider crate? Would we be able to test Pull
> > > > > >>> Requests on
> > > > > >>> > Iceberg Rust with sqllogictests?
> > > > > >>> >
> > > > > >>> > 3) @kevinqliu is familiar with the Python bindings in Iceberg
> > > Rust,
> > > > > >>> and I
> > > > > >>> > believe that DataFusion dependency might be removed as well,
> > but
> > > that
> > > > > >>> is
> > > > > >>> > not the core focus of this conversation. He has also proposed
> > > removing
> > > > > >>> that
> > > > > >>> > and using the DataFusion Python bindings directly.
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > I'm sure I'm forgetting things, but this email is long enough.
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > Thanks everyone for the discussion thus far, and looking
> > forward
> > > to
> > > > > >>> more
> > > > > >>> > input on this.
> > > > > >>> >
> > > > > >>> >
> > > > > >>> > -Matt
> > > > > >>> >
> > > > > >>> >
> > > > > >>> >
> > > > > >>> >
> > > > > >>>
> > > > > >>
> > > > >
> > > >
> > > > Xuanwo
> > > >
> > > > https://xuanwo.io/
> > > >
> > >
> >
> 

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to