+1 (non-binding) On Fri, Sep 11, 2026 at 12:40 PM Jerry Peng <[email protected]> wrote:
> +1 (non-binding) > > > On Fri, Sep 11, 2026 at 9:03 AM Stevo Mitrić via dev <[email protected]> > wrote: > >> +1 (non-binding) >> >> On Fri, Sep 11, 2026 at 8:02 AM Andreas Neumann <[email protected]> wrote: >> >>> +1 (non-binding) >>> >>> On Thu, Sep 10, 2026 at 7:34 PM Ye Xianjin <[email protected]> wrote: >>> >>>> +1 >>>> Sent from my iPhone >>>> >>>> On Sep 11, 2026, at 10:19 AM, Yuming Wang <[email protected]> wrote: >>>> >>>> >>>> +1 >>>> >>>> On Fri, Sep 11, 2026 at 7:07 AM Holden Karau <[email protected]> >>>> wrote: >>>> >>>>> +1 >>>>> >>>>> >>>>> Twitter: https://twitter.com/holdenkarau >>>>> Fight Health Insurance: https://www.fighthealthinsurance.com/ >>>>> <https://www.fighthealthinsurance.com/?q=hk_email> >>>>> Books (Learning Spark, High Performance Spark, etc.): >>>>> https://amzn.to/2MaRAG9 <https://amzn.to/2MaRAG9> >>>>> YouTube Live Streams: https://www.youtube.com/user/holdenkarau >>>>> Pronouns: she/her >>>>> >>>>> On Fri, Sep 11, 2026 at 4:08 AM huaxin gao <[email protected]> >>>>> wrote: >>>>> >>>>>> +1 >>>>>> >>>>>> On Thu, Sep 10, 2026 at 12:07 PM karuppayya <[email protected]> >>>>>> wrote: >>>>>> >>>>>>> +1 (non-binding) >>>>>>> >>>>>>> On Thu, Sep 10, 2026 at 11:30 AM John Zhuge <[email protected]> >>>>>>> wrote: >>>>>>> >>>>>>>> +1 (non binding) >>>>>>>> >>>>>>>> Very useful. >>>>>>>> >>>>>>>> On Thu, Sep 10, 2026 at 11:11 AM Szehon Ho <[email protected]> >>>>>>>> wrote: >>>>>>>> >>>>>>>>> +1 (non binding) >>>>>>>>> >>>>>>>>> Thanks and excited for the File type >>>>>>>>> Szehon >>>>>>>>> >>>>>>>>> On Wed, Sep 9, 2026 at 11:53 AM Daniel Tenedorio < >>>>>>>>> [email protected]> wrote: >>>>>>>>> >>>>>>>>>> +1 (non-binding) from me as well. This will make Spark a stronger >>>>>>>>>> engine for processing large values, such as large images or video >>>>>>>>>> clips. >>>>>>>>>> Avoiding materializing the values at shuffle boundaries or other >>>>>>>>>> points >>>>>>>>>> until we actually need to consume the values can make pipelines >>>>>>>>>> stable and >>>>>>>>>> performant. >>>>>>>>>> >>>>>>>>>> On 2026/09/08 22:14:35 Yicong Huang wrote: >>>>>>>>>> > +1 (non-binding) >>>>>>>>>> > >>>>>>>>>> > I like the idea, especially for UDF to understand the FILE >>>>>>>>>> semantic. >>>>>>>>>> > >>>>>>>>>> > An extended idea is if it makes sense to even support a >>>>>>>>>> folder/directory as a collection of FILEs. I see many use cases have >>>>>>>>>> a >>>>>>>>>> dataset (e.g., images) in a folder, and if spark can understand >>>>>>>>>> that's a >>>>>>>>>> collection of FILEs it would be great to handle their life cycles. >>>>>>>>>> > >>>>>>>>>> > Best, >>>>>>>>>> > Yicong >>>>>>>>>> > >>>>>>>>>> > >>>>>>>>>> > On 2026/09/08 22:03:52 Gengliang Wang wrote: >>>>>>>>>> > > +1 >>>>>>>>>> > > >>>>>>>>>> > > On Tue, Sep 8, 2026 at 2:59 PM Hyukjin Kwon < >>>>>>>>>> [email protected]> wrote: >>>>>>>>>> > > >>>>>>>>>> > > > +1 >>>>>>>>>> > > > >>>>>>>>>> > > > On 2026/09/08 20:34:38 Xiao Li wrote: >>>>>>>>>> > > > > +1 from me. This fills a real gap in Spark. Modeling >>>>>>>>>> files as either >>>>>>>>>> > > > binary >>>>>>>>>> > > > > blobs or opaque paths is awkward, especially as >>>>>>>>>> unstructured data >>>>>>>>>> > > > workloads >>>>>>>>>> > > > > grow. >>>>>>>>>> > > > > >>>>>>>>>> > > > > A first-class FILE type, aligned with Parquet and with >>>>>>>>>> lazy content >>>>>>>>>> > > > > loading, feels like the right abstraction. I support >>>>>>>>>> moving this forward. >>>>>>>>>> > > > > >>>>>>>>>> > > > > Xiao >>>>>>>>>> > > > > >>>>>>>>>> > > > > Burak Yavuz <[email protected]> 于2026年9月8日周二 12:49写道: >>>>>>>>>> > > > > >>>>>>>>>> > > > > > I'll kick off the vote with a +1 (non-binding) >>>>>>>>>> > > > > > >>>>>>>>>> > > > > > Thanks, >>>>>>>>>> > > > > > Burak >>>>>>>>>> > > > > > >>>>>>>>>> > > > > > On Tue, Sep 8, 2026 at 3:47 PM Burak Yavuz < >>>>>>>>>> [email protected]> wrote: >>>>>>>>>> > > > > > >>>>>>>>>> > > > > >> Hi Spark devs, >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> I would like to start a vote on introducing FileType >>>>>>>>>> for handling >>>>>>>>>> > > > > >> unstructured data. >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> The SPIP document: >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> >>>>>>>>>> > > > >>>>>>>>>> https://docs.google.com/document/d/1pPof896ZwcZ-2Yn-YhC4TbyYWn1umiQJykGDkAxzCWc/edit?tab=t.0#heading=h.m1700lw4wsoj >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> Discussion thread: >>>>>>>>>> > > > > >> >>>>>>>>>> https://lists.apache.org/thread/6f83qcpj1ox40jfxtottqlhjbhhxwmhf >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> JIRA Ticket: >>>>>>>>>> > > > > >> https://issues.apache.org/jira/browse/SPARK-59132 >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> The vote will be open for at least 72 hours, and >>>>>>>>>> passes if a majority >>>>>>>>>> > > > +1 >>>>>>>>>> > > > > >> PMC >>>>>>>>>> > > > > >> votes are cast, with a minimum of 3 +1 votes. >>>>>>>>>> > > > > >> Please vote: >>>>>>>>>> > > > > >> [ ] +1: Accept the proposal as an official SPIP >>>>>>>>>> > > > > >> [ ] +0 >>>>>>>>>> > > > > >> [ ] -1: I don't think this is a good idea because ... >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> Best regards, >>>>>>>>>> > > > > >> Burak Yavuz >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >> >>>>>>>>>> > > > > >>>>>>>>>> > > > >>>>>>>>>> > > > >>>>>>>>>> --------------------------------------------------------------------- >>>>>>>>>> > > > To unsubscribe e-mail: [email protected] >>>>>>>>>> > > > >>>>>>>>>> > > > >>>>>>>>>> > > >>>>>>>>>> > >>>>>>>>>> > >>>>>>>>>> --------------------------------------------------------------------- >>>>>>>>>> > To unsubscribe e-mail: [email protected] >>>>>>>>>> > >>>>>>>>>> > >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> --------------------------------------------------------------------- >>>>>>>>>> To unsubscribe e-mail: [email protected] >>>>>>>>>> >>>>>>>>>> >>>>>>>> >>>>>>>> -- >>>>>>>> John Zhuge >>>>>>>> >>>>>>> -- Regards Rishab Joshi
