I actually support this. I got some questions from several ppl in the
community who want to contribute to this.

On Wed, 19 Aug 2026 at 08:46, Allison Wang <[email protected]> wrote:

> Hi all,
>
> I would like to discuss whether Apache Spark should provide an experimental,
> community-maintained home for the ecosystem around the PySpark Data
> Source API
> <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html>
> :
>
>
> Proposed repository: apache/spark-python-datasources
>
> Existing implementation: allisonwang-db/pyspark-data-sources
> <https://github.com/allisonwang-db/pyspark-data-sources>
>
> The intent would not be to make these data sources part of Spark core.
> Instead, the repository would provide an Apache-governed place where the
> community can collaborate on reusable implementations of the public Python
> Data Source API, share practical examples, and grow the ecosystem around
> the API without expanding Spark core itself.
>
> The existing project can serve as the initial contribution. It contains
> batch and streaming readers and writers built with the public PySpark Data
> Source API, covering a range of external systems and use cases.
>
> I propose starting with a deliberately lightweight model:
>
>    -
>
>    The repository would be experimental and community-supported.
>    -
>
>    It would focus specifically on implementations built on the public PySpark
>    Data Source API.
>    
> <https://spark.apache.org/docs/latest/api/python/tutorial/sql/python_data_source.html>
>    -
>
>    Code in this repository would remain separate from Spark core and
>    would not carry the same compatibility or support guarantees as Spark
>    itself.
>    -
>
>    New contributions would go through community review, with
>    maintainability, dependencies, licensing, testing, and security taken into
>    consideration.
>    -
>
>    Implementations that become unmaintained or no longer meet the
>    repository's requirements could be deprecated or removed through the normal
>    community process.
>    -
>
>    The repository would be governed by the Apache Spark community and
>    follow ASF policies.
>
> The existing project is currently published on PyPI as
> pyspark-data-sources using the pyspark_datasources import namespace. For
> continuity, I would prefer to retain those names if they are compatible
> with ASF release and branding requirements, but the package naming is not
> essential to this proposal. I am willing to help maintain the repository,
> review contributions, and support the release process.
>
> The main question I would like feedback on is whether the Spark community
> thinks it is useful to provide this kind of lightweight, experimental
> home for extensions built on a public Spark API, while keeping those
> integrations explicitly outside Spark core.
>
> If the community supports this proposal, I will work with the Spark PMC on
> the required JIRA and ASF IP-clearance steps, move the approved code to the
> Apache repository, update the package metadata, and transfer PyPI
> publishing to an ASF-controlled release process. Existing PyPI releases and
> installation commands would remain unchanged.
>
> I would appreciate any feedback on this proposal.
>
> Thanks,
>
> Allison
>
>

Reply via email to