Just to be clear, I mean "broadly similar” as in they are both 
community-maintained libraries for Spark. Allison’s proposal is, of course, 
much more focused! I would just like to raise the point of comparison since I 
think it’s relevant.


> On Aug 18, 2026, at 10:51 PM, Nicholas Chammas <[email protected]> 
> wrote:
> 
> We had something broadly similar from ~10 years ago, but not limited to data 
> sources: https://spark-packages.org/
> 
> My impression of Spark Packages is that it is inactive and unmaintained. If 
> so, it might be useful to have a brief post mortem as a community to 
> understand why it didn’t stand the test of time, and how to avoid that fate 
> for this new proposal.
> 
> 
>> On Aug 18, 2026, at 7:45 PM, Allison Wang <[email protected]> wrote:
>> 
>> Hi all,
>> 
>> I would like to discuss whether Apache Spark should provide an experimental, 
>> community-maintained home for the ecosystem around the PySpark Data Source 
>> API 
>> <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html>:
>> 
>> Proposed repository: apache/spark-python-datasources
>> Existing implementation: allisonwang-db/pyspark-data-sources 
>> <https://github.com/allisonwang-db/pyspark-data-sources>
>> The intent would not be to make these data sources part of Spark core. 
>> Instead, the repository would provide an Apache-governed place where the 
>> community can collaborate on reusable implementations of the public Python 
>> Data Source API, share practical examples, and grow the ecosystem around the 
>> API without expanding Spark core itself.
>> 
>> The existing project can serve as the initial contribution. It contains 
>> batch and streaming readers and writers built with the public PySpark Data 
>> Source API, covering a range of external systems and use cases.
>> 
>> I propose starting with a deliberately lightweight model:
>> 
>> The repository would be experimental and community-supported.
>> It would focus specifically on implementations built on the public PySpark 
>> Data Source API. 
>> <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html>
>> Code in this repository would remain separate from Spark core and would not 
>> carry the same compatibility or support guarantees as Spark itself.
>> New contributions would go through community review, with maintainability, 
>> dependencies, licensing, testing, and security taken into consideration.
>> Implementations that become unmaintained or no longer meet the repository's 
>> requirements could be deprecated or removed through the normal community 
>> process.
>> The repository would be governed by the Apache Spark community and follow 
>> ASF policies.
>> 
>> The existing project is currently published on PyPI as pyspark-data-sources 
>> using the pyspark_datasources import namespace. For continuity, I would 
>> prefer to retain those names if they are compatible with ASF release and 
>> branding requirements, but the package naming is not essential to this 
>> proposal. I am willing to help maintain the repository, review 
>> contributions, and support the release process.
>> 
>> The main question I would like feedback on is whether the Spark community 
>> thinks it is useful to provide this kind of lightweight, experimental home 
>> for extensions built on a public Spark API, while keeping those integrations 
>> explicitly outside Spark core.
>> 
>> If the community supports this proposal, I will work with the Spark PMC on 
>> the required JIRA and ASF IP-clearance steps, move the approved code to the 
>> Apache repository, update the package metadata, and transfer PyPI publishing 
>> to an ASF-controlled release process. Existing PyPI releases and 
>> installation commands would remain unchanged.
>> 
>> I would appreciate any feedback on this proposal.
>> 
>> Thanks,
>> Allison
>> 
> 

Reply via email to