+1
Twitter: https://twitter.com/holdenkarau Fight Health Insurance: https://www.fighthealthinsurance.com/ <https://www.fighthealthinsurance.com/?q=hk_email> Books (Learning Spark, High Performance Spark, etc.): https://amzn.to/2MaRAG9 <https://amzn.to/2MaRAG9> YouTube Live Streams: https://www.youtube.com/user/holdenkarau Pronouns: she/her On Fri, Sep 11, 2026 at 4:08 AM huaxin gao <[email protected]> wrote: > +1 > > On Thu, Sep 10, 2026 at 12:07 PM karuppayya <[email protected]> > wrote: > >> +1 (non-binding) >> >> On Thu, Sep 10, 2026 at 11:30 AM John Zhuge <[email protected]> wrote: >> >>> +1 (non binding) >>> >>> Very useful. >>> >>> On Thu, Sep 10, 2026 at 11:11 AM Szehon Ho <[email protected]> >>> wrote: >>> >>>> +1 (non binding) >>>> >>>> Thanks and excited for the File type >>>> Szehon >>>> >>>> On Wed, Sep 9, 2026 at 11:53 AM Daniel Tenedorio <[email protected]> >>>> wrote: >>>> >>>>> +1 (non-binding) from me as well. This will make Spark a stronger >>>>> engine for processing large values, such as large images or video clips. >>>>> Avoiding materializing the values at shuffle boundaries or other points >>>>> until we actually need to consume the values can make pipelines stable and >>>>> performant. >>>>> >>>>> On 2026/09/08 22:14:35 Yicong Huang wrote: >>>>> > +1 (non-binding) >>>>> > >>>>> > I like the idea, especially for UDF to understand the FILE semantic. >>>>> > >>>>> > An extended idea is if it makes sense to even support a >>>>> folder/directory as a collection of FILEs. I see many use cases have a >>>>> dataset (e.g., images) in a folder, and if spark can understand that's a >>>>> collection of FILEs it would be great to handle their life cycles. >>>>> > >>>>> > Best, >>>>> > Yicong >>>>> > >>>>> > >>>>> > On 2026/09/08 22:03:52 Gengliang Wang wrote: >>>>> > > +1 >>>>> > > >>>>> > > On Tue, Sep 8, 2026 at 2:59 PM Hyukjin Kwon <[email protected]> >>>>> wrote: >>>>> > > >>>>> > > > +1 >>>>> > > > >>>>> > > > On 2026/09/08 20:34:38 Xiao Li wrote: >>>>> > > > > +1 from me. This fills a real gap in Spark. Modeling files as >>>>> either >>>>> > > > binary >>>>> > > > > blobs or opaque paths is awkward, especially as unstructured >>>>> data >>>>> > > > workloads >>>>> > > > > grow. >>>>> > > > > >>>>> > > > > A first-class FILE type, aligned with Parquet and with lazy >>>>> content >>>>> > > > > loading, feels like the right abstraction. I support moving >>>>> this forward. >>>>> > > > > >>>>> > > > > Xiao >>>>> > > > > >>>>> > > > > Burak Yavuz <[email protected]> 于2026年9月8日周二 12:49写道: >>>>> > > > > >>>>> > > > > > I'll kick off the vote with a +1 (non-binding) >>>>> > > > > > >>>>> > > > > > Thanks, >>>>> > > > > > Burak >>>>> > > > > > >>>>> > > > > > On Tue, Sep 8, 2026 at 3:47 PM Burak Yavuz <[email protected]> >>>>> wrote: >>>>> > > > > > >>>>> > > > > >> Hi Spark devs, >>>>> > > > > >> >>>>> > > > > >> I would like to start a vote on introducing FileType for >>>>> handling >>>>> > > > > >> unstructured data. >>>>> > > > > >> >>>>> > > > > >> The SPIP document: >>>>> > > > > >> >>>>> > > > > >> >>>>> > > > >>>>> https://docs.google.com/document/d/1pPof896ZwcZ-2Yn-YhC4TbyYWn1umiQJykGDkAxzCWc/edit?tab=t.0#heading=h.m1700lw4wsoj >>>>> > > > > >> >>>>> > > > > >> Discussion thread: >>>>> > > > > >> >>>>> https://lists.apache.org/thread/6f83qcpj1ox40jfxtottqlhjbhhxwmhf >>>>> > > > > >> >>>>> > > > > >> JIRA Ticket: >>>>> > > > > >> https://issues.apache.org/jira/browse/SPARK-59132 >>>>> > > > > >> >>>>> > > > > >> The vote will be open for at least 72 hours, and passes if >>>>> a majority >>>>> > > > +1 >>>>> > > > > >> PMC >>>>> > > > > >> votes are cast, with a minimum of 3 +1 votes. >>>>> > > > > >> Please vote: >>>>> > > > > >> [ ] +1: Accept the proposal as an official SPIP >>>>> > > > > >> [ ] +0 >>>>> > > > > >> [ ] -1: I don't think this is a good idea because ... >>>>> > > > > >> >>>>> > > > > >> Best regards, >>>>> > > > > >> Burak Yavuz >>>>> > > > > >> >>>>> > > > > >> >>>>> > > > > >> >>>>> > > > > >>>>> > > > >>>>> > > > >>>>> --------------------------------------------------------------------- >>>>> > > > To unsubscribe e-mail: [email protected] >>>>> > > > >>>>> > > > >>>>> > > >>>>> > >>>>> > --------------------------------------------------------------------- >>>>> > To unsubscribe e-mail: [email protected] >>>>> > >>>>> > >>>>> >>>>> --------------------------------------------------------------------- >>>>> To unsubscribe e-mail: [email protected] >>>>> >>>>> >>> >>> -- >>> John Zhuge >>> >>
