Just to be clear, I mean "broadly similar” as in they are both community-maintained libraries for Spark. Allison’s proposal is, of course, much more focused! I would just like to raise the point of comparison since I think it’s relevant.
> On Aug 18, 2026, at 10:51 PM, Nicholas Chammas <[email protected]> > wrote: > > We had something broadly similar from ~10 years ago, but not limited to data > sources: https://spark-packages.org/ > > My impression of Spark Packages is that it is inactive and unmaintained. If > so, it might be useful to have a brief post mortem as a community to > understand why it didn’t stand the test of time, and how to avoid that fate > for this new proposal. > > >> On Aug 18, 2026, at 7:45 PM, Allison Wang <[email protected]> wrote: >> >> Hi all, >> >> I would like to discuss whether Apache Spark should provide an experimental, >> community-maintained home for the ecosystem around the PySpark Data Source >> API >> <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html>: >> >> Proposed repository: apache/spark-python-datasources >> Existing implementation: allisonwang-db/pyspark-data-sources >> <https://github.com/allisonwang-db/pyspark-data-sources> >> The intent would not be to make these data sources part of Spark core. >> Instead, the repository would provide an Apache-governed place where the >> community can collaborate on reusable implementations of the public Python >> Data Source API, share practical examples, and grow the ecosystem around the >> API without expanding Spark core itself. >> >> The existing project can serve as the initial contribution. It contains >> batch and streaming readers and writers built with the public PySpark Data >> Source API, covering a range of external systems and use cases. >> >> I propose starting with a deliberately lightweight model: >> >> The repository would be experimental and community-supported. >> It would focus specifically on implementations built on the public PySpark >> Data Source API. >> <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html> >> Code in this repository would remain separate from Spark core and would not >> carry the same compatibility or support guarantees as Spark itself. >> New contributions would go through community review, with maintainability, >> dependencies, licensing, testing, and security taken into consideration. >> Implementations that become unmaintained or no longer meet the repository's >> requirements could be deprecated or removed through the normal community >> process. >> The repository would be governed by the Apache Spark community and follow >> ASF policies. >> >> The existing project is currently published on PyPI as pyspark-data-sources >> using the pyspark_datasources import namespace. For continuity, I would >> prefer to retain those names if they are compatible with ASF release and >> branding requirements, but the package naming is not essential to this >> proposal. I am willing to help maintain the repository, review >> contributions, and support the release process. >> >> The main question I would like feedback on is whether the Spark community >> thinks it is useful to provide this kind of lightweight, experimental home >> for extensions built on a public Spark API, while keeping those integrations >> explicitly outside Spark core. >> >> If the community supports this proposal, I will work with the Spark PMC on >> the required JIRA and ASF IP-clearance steps, move the approved code to the >> Apache repository, update the package metadata, and transfer PyPI publishing >> to an ASF-controlled release process. Existing PyPI releases and >> installation commands would remain unchanged. >> >> I would appreciate any feedback on this proposal. >> >> Thanks, >> Allison >> >
