We had something broadly similar from ~10 years ago, but not limited to data sources: https://spark-packages.org/
My impression of Spark Packages is that it is inactive and unmaintained. If so, it might be useful to have a brief post mortem as a community to understand why it didn’t stand the test of time, and how to avoid that fate for this new proposal. > On Aug 18, 2026, at 7:45 PM, Allison Wang <[email protected]> wrote: > > Hi all, > > I would like to discuss whether Apache Spark should provide an experimental, > community-maintained home for the ecosystem around the PySpark Data Source > API > <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html>: > > Proposed repository: apache/spark-python-datasources > Existing implementation: allisonwang-db/pyspark-data-sources > <https://github.com/allisonwang-db/pyspark-data-sources> > The intent would not be to make these data sources part of Spark core. > Instead, the repository would provide an Apache-governed place where the > community can collaborate on reusable implementations of the public Python > Data Source API, share practical examples, and grow the ecosystem around the > API without expanding Spark core itself. > > The existing project can serve as the initial contribution. It contains batch > and streaming readers and writers built with the public PySpark Data Source > API, covering a range of external systems and use cases. > > I propose starting with a deliberately lightweight model: > > The repository would be experimental and community-supported. > It would focus specifically on implementations built on the public PySpark > Data Source API. > <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html> > Code in this repository would remain separate from Spark core and would not > carry the same compatibility or support guarantees as Spark itself. > New contributions would go through community review, with maintainability, > dependencies, licensing, testing, and security taken into consideration. > Implementations that become unmaintained or no longer meet the repository's > requirements could be deprecated or removed through the normal community > process. > The repository would be governed by the Apache Spark community and follow ASF > policies. > > The existing project is currently published on PyPI as pyspark-data-sources > using the pyspark_datasources import namespace. For continuity, I would > prefer to retain those names if they are compatible with ASF release and > branding requirements, but the package naming is not essential to this > proposal. I am willing to help maintain the repository, review contributions, > and support the release process. > > The main question I would like feedback on is whether the Spark community > thinks it is useful to provide this kind of lightweight, experimental home > for extensions built on a public Spark API, while keeping those integrations > explicitly outside Spark core. > > If the community supports this proposal, I will work with the Spark PMC on > the required JIRA and ASF IP-clearance steps, move the approved code to the > Apache repository, update the package metadata, and transfer PyPI publishing > to an ASF-controlled release process. Existing PyPI releases and installation > commands would remain unchanged. > > I would appreciate any feedback on this proposal. > > Thanks, > Allison >
