Hi Holden,

Thanks for the feedback. These are valid concerns. I will withdraw
this vote for now and follow up with a revised proposal that addresses
these concerns.

Thanks,
Allison

On Fri, Aug 28, 2026 at 12:14 AM Holden Karau <[email protected]> wrote:
>
> I don’t think is a good idea *yet*
>
> My core hesitation stems from the fact that it is taking on a bunch of 
> connectors that the existing committers don’t necessarily have context 
> around. Historically connectors maintained by the data sources or sink 
> communities / companies themselves have done better than those inside of an 
> amalgamation or inside the project.
>
> I’m not convinced that AI meaningfully changes the problems we’ve seen with 
> in tree attempts.
>
> There’s not a very clear “this is our release plan” yet or what the quality 
> bar would be just that it’s “different.”
>
> Also looking at the connector package it offers datasources which have their 
> own existing out of tree Spark version. How would those developer feel / how 
> is our version better / will this cause user confusion?
>
> Additionally I’m a little cautious about the “based on” given the direct 
> databricks reference for comparability in the readme we’d definitely need to 
> clean that up.
>
> For those reasons I’m -1 on this iteration of the proposal.
>
>
> Twitter: https://twitter.com/holdenkarau
> Fight Health Insurance: https://www.fighthealthinsurance.com/
> Books (Learning Spark, High Performance Spark, etc.): https://amzn.to/2MaRAG9
> YouTube Live Streams: https://www.youtube.com/user/holdenkarau
> Pronouns: she/her
>
> On Thu, Aug 27, 2026 at 5:42 PM Hyukjin Kwon <[email protected]> wrote:
>>
>> +1
>>
>> On Fri, 28 Aug 2026 at 03:27, Uroš Bojanić <[email protected]> wrote:
>>>
>>> +1 (non-binding)
>>>
>>> On 2026/08/27 03:36:05 Allison Wang wrote:
>>> > Hi all,
>>> >
>>> > I'd like to open the vote on the proposal to introduce a new Apache Spark
>>> > repository for community-contributed Python data sources:
>>> > *apache/spark-python-data-sources* based on the existing project
>>> > https://github.com/allisonwang-db/pyspark-data-sources
>>> >
>>> > DISCUSS thread:
>>> > https://lists.apache.org/thread/0yqzbr8pzfnfomqq1zdgzcqltnrc6s1o
>>> >
>>> > The vote is open for at least the next 72 hours.
>>> >
>>> > [ ] +1: Accept the proposal
>>> > [ ] +0
>>> > [ ] -1: I don't think this is a good idea because...
>>> >
>>> > Thanks,
>>> > Allison
>>> >
>>>
>>> ---------------------------------------------------------------------
>>> To unsubscribe e-mail: [email protected]
>>>

---------------------------------------------------------------------
To unsubscribe e-mail: [email protected]

Reply via email to