Hi Matt,

Thanks for raising this! I generally agree that we can move the datafusion
integration to a separate repo if that helps more DataFusion experts to
work on the integrations

On the three considerations:
1) Apache governance:
I think this is my biggest concern so far. This could not only affect
contributors, but also whether the downstream users could continue to use
the integration.
Even if the code remains Apache-licensed, governance, release of artifacts,
and contribution policies can matter for adoption.
We should explore more about the option to keep it under an Apache-governed
repository.

2) Where do tests live/how tests should be maintained?

I think this is closely related to the governance question. To me, the
broader question is: *how do we make it easy for people from both the
Iceberg Rust and DataFusion communities to maintain the integration?*

If we move it to a non-Apache repository, I worry that it could eventually
look somewhat like the current Iceberg Java <> Trino integration: the
integration primarily lives on the Trino side and is therefore mostly
maintained by people who are already deeply involved in Trino.

There is an important difference here, though. For Iceberg Java, Spark is
arguably the primary engine integration and has a large Iceberg contributor
base around it, so having the Trino integration maintained more
independently is relatively natural. In Iceberg Rust today, DataFusion has
a much more central role. It is by far the most mature engine integration
in the project, is used by our SQLLogicTest infrastructure, and many users
building on Iceberg Rust are also building on DataFusion. The overlap
between the two communities is therefore much larger.

Because of that, I would prefer that extracting the integration does not
turn it into something that is effectively owned only by the DataFusion
side. Ideally, both Iceberg Rust contributors and DataFusion contributors
should be able to review changes, maintain compatibility, and evolve the
integration together.

3) Dropping DataFusion dependencies from pyiceberg-core binding: I think it
makes sense to make things simpler and have left a comment on the github
issue with more detailed thoughts:
https://github.com/apache/iceberg-rust/issues/3036

Best,
Shawn

On Fri, Aug 21, 2026 at 11:23 AM Matt Butrovich <[email protected]>
wrote:

> Hello all,
>
>
> I've never tried emailing two different project lists at once, but it was
> suggested that I do so to try to track the conversation across both
> communities. We'll see how this threads on the mailing lists.
>
>
> There has been a GitHub Discussion
> https://github.com/apache/iceberg-rust/discussions/2992#discussioncomment-18082533,
> a GitHub Issue https://github.com/apache/iceberg-rust/issues/3029, and
> it's been a long topic of conversation in the past two weeks in both the
> DataFusion and Iceberg Rust Community Calls to discuss moving the
> DataFusion integration from Iceberg Rust to a separate repository. I will
> try to summarize some of the major points as I understand them, but the
> conversations are the ground truth and please feel free to correct me here.
>
>
> The DataFusion TableProvider integration in Iceberg Rust makes DataFusion
> a dependency for Iceberg Rust. The integration exists for multiple reasons:
>
> 1) an engine to execute Iceberg Rust's corpus of sqllogictest files
>
> 2) a TableProvider integration for DataFusion users to interact with
> Iceberg tables
>
>
> Some of the motivations to break out the integration:
>
> 1) There have been a number of issues and pull requests against the
> DataFusion TableProvider in Iceberg Rust as users want to add more
> features, and they often go stale. I don't believe there are many
> committers/PMC members familiar with or using the DataFusion integration.
>
> 2) Iceberg Rust would like to stay as engine-agnostic as possible. A
> recent DataFusion Ballista integration was declined for this reason
> https://github.com/apache/iceberg-rust/pull/2613.
>
> 3) Other projects that rely on both Iceberg Rust and DataFusion (e.g.,
> Comet) are blocked by Iceberg Rust upgrading its DataFusion and Arrow
> dependencies before they can upgrade.
>
>
> Some considerations for both communities:
>
> 1) Where would this DataFusion TableProvider live? Most specifically,
> would it be an Apache-governed project? As folks like @andygrove point out,
> this can affect whether some community members could contribute to it.
> There is a datafusion-contrib org for DataFusion-related projects to have
> visibility but no Apache governance, but there may be options to put it
> under an Apache repository.
>
> 2) How would Iceberg Rust continue to run sqllogictests for regression
> testing? Does this live in a different repository that depends on this new
> Iceberg Rust TableProvider crate? Would we be able to test Pull Requests on
> Iceberg Rust with sqllogictests?
>
> 3) @kevinqliu is familiar with the Python bindings in Iceberg Rust, and I
> believe that DataFusion dependency might be removed as well, but that is
> not the core focus of this conversation. He has also proposed removing that
> and using the DataFusion Python bindings directly.
>
>
> I'm sure I'm forgetting things, but this email is long enough.
>
>
> Thanks everyone for the discussion thus far, and looking forward to more
> input on this.
>
>
> -Matt
>
>
>
>

Reply via email to