What do we think about adding version discovery / capability discovery so
we can do more intelligent routing? Probably useful in layer 2


Twitter: https://twitter.com/holdenkarau
Fight Health Insurance: https://www.fighthealthinsurance.com/
<https://www.fighthealthinsurance.com/?q=hk_email>
Books (Learning Spark, High Performance Spark, etc.):
https://amzn.to/2MaRAG9  <https://amzn.to/2MaRAG9>
YouTube Live Streams: https://www.youtube.com/user/holdenkarau
Pronouns: she/her

On Sat, Aug 1, 2026 at 4:54 AM Adam Binford <[email protected]> wrote:

> I see the README in the project actually lists my proxy as another
> existing implementation, thanks for the call out. It's also Rust based for
> performance, but proxies raw HTTP/2 traffic instead of parsing and
> forwarding gRPC. As one of the people who has built a custom gateway/proxy
> like this, there's a few use cases it doesn't meet, mostly for cases of
> fine-grained data access controls per user, not necessarily per "tenant". I
> know on-prem HDFS/Kerberos based users are probably a dying breed, so I'm
> not sure how many other people actually need that functionality, but in
> this environment a pre-created Spark Driver that's shared across multiple
> users simply doesn't work. I'm a little curious how this works for other
> existing use cases but for us:
> - Authorization to data is tied to the user running the Spark job, so each
> user needs its own job/driver
> - Spark drivers can't be preallocated for every user, the users need to be
> able to opt in to creating them
> - Even if users could share drivers, we would constantly deal with dumb
> users killing other users sessions with driver OOM or other poor practices,
> each user having its own driver at least limits the blast radius.
>
> That's why one of the main purposes of my implementation is a REST API for
> creating new Spark jobs, which then maps everything together with static
> tokens at the HTTP/2 level, so no gRPC traffic needs to be parsed. This
> might just be a custom enough use case for my environment, but I'm curious
> if anyone else would need similar functionality to actually make this
> useful.
>
> To be clear though, my project is very much a pet project that is not
> heavily battle tested. I was also thinking if similar functionality was
> ever upstreamed it would likely be in Scala, though Rust is the perfect
> language for such a thing, and one of the main reasons you call out is
> hyper's support for trailing headers.
>
> I do think it's useful to have some kind of official implementation of
> this, as there is a lot of extra work required to actually make Spark
> Connect useful in real environments.
>
> Adam
>
> On Sat, Aug 1, 2026 at 6:32 AM Mark Hamstra <[email protected]> wrote:
>
>> I can't say that the proposal to use Rust or any other non-JVM
>> language for something that feels like a new core component of Spark
>> sits easy with me, That is not to say that I am firmly opposed or
>> can't be persuaded, but just that I am now uneasy. The arguments in
>> the SPIP for using Rust don't strike me as particularly strong, being
>> closer to "I/we like Rust" rather than a more persuasive
>> compare-and-contrast of the relative merits for this component written
>> and maintained in Rust, Scala or Java. The idea that the basic design
>> of this gateway should be independent enough that it could conceivably
>> be written in any of several languages without the end users caring
>> much or at all is, I think, a good one, but that on its own doesn't
>> persuade me that it shouldn't be written in the same version of Scala
>> as the rest of core Spark, or in Scala 3, which has considerable
>> advantages over Scala 2 for those components that are independent
>> enough from the Scala 2 Spark core.
>>
>> Also, this new gateway does not strike me as analogous to
>> spark-connect-swift or spark-connect-go or spark-connect-rust or some
>> of the nascent talk about something like spark-connect-scala3 since
>> those are all efforts to provide functionality for particular
>> language-using subsets of the community, not for more or less
>> everyone.
>>
>> No conclusions from me yet, just wanting to see more discussion before
>> deciding.
>>
>> On Fri, Jul 31, 2026 at 11:54 PM vaquar khan <[email protected]>
>> wrote:
>> >
>> > +1
>> >
>> > A community-maintained implementation already exists for this
>> functionality. I have added a few comments to the SIP document outlining
>> how this prior art relates to the current proposal.
>> >
>> > Regards,
>> > Viquar Khan
>> >
>> > On Fri, 31 Jul 2026 at 23:37, Dongjoon Hyun <[email protected]>
>> wrote:
>> >>
>> >> For the record, Apache Spark community currently maintains seven
>> GitHub repositories like the following:
>> >>
>> >> - https://github.com/apache/spark
>> (Last Commit: Today)
>> >> - https://github.com/apache/spark-connect-swift             (Last
>> Commit: Today)
>> >> - https://github.com/apache/spark-kubernetes-operator  (Last Commit:
>> 3 days ago)
>> >> - https://github.com/apache/spark-website                       (Last
>> Commit 5 days ago)
>> >> - https://github.com/apache/spark-docker                        (Last
>> Commit: 1 week ago)
>> >> - https://github.com/apache/spark-connect-rust              (Last
>> Commit: 2 months ago)
>> >> - https://github.com/apache/spark-connect-go                (Last
>> Commit: 11 months ago)
>> >>
>> >> Since there is no bug-free software, I hope the new proposed
>> repository builds a new sub-community by being one of active Apache Spark
>> repositories instead of one-time snapshot drop-off.
>> >>
>> >> - https://github.com/apache/spark-connect-gateway
>> >>
>> >> Sincerely,
>> >> Dongjoon.
>> >>
>> >> On 2026/08/01 04:30:05 John Zhuge wrote:
>> >> > +1 Thanks.
>> >> >
>> >> > On Fri, Jul 31, 2026 at 8:41 PM Dongjoon Hyun <[email protected]>
>> wrote:
>> >> >
>> >> > > Thank you, Liang-Chi.
>> >> > >
>> >> > > Since it's written in Rust like "apache/spark-connect-rust",
>> >> > > "apache/spark-connect-gateway" proposal sounds good to me.
>> >> > >
>> >> > > Where is the reference implementation? I didn't find any GitHub
>> link in
>> >> > > SPARK-58455 and SPIP.
>> >> > >
>> >> > > Dongjoon.
>> >> > >
>> >> > > On 2026/07/31 22:49:44 Chao Sun wrote:
>> >> > > > +1
>> >> > > >
>> >> > > > On Fri, Jul 31, 2026 at 1:26 PM DB Tsai <[email protected]>
>> wrote:
>> >> > > >
>> >> > > > > +1
>> >> > > > >
>> >> > > > > An open-source reference gateway that makes it easier to run
>> Spark
>> >> > > Connect
>> >> > > > > at scale with session affinity, authentication, high
>> availability,
>> >> > > > > multi-tenancy, rate limiting, and audit logging built in is
>> great for
>> >> > > > > production adoption of Spark Connect.
>> >> > > > >
>> >> > > > >
>> >> > > > > DB Tsai  |  https://www.dbtsai.com/  |  PGP 0x9FB9FAA3
>> >> > > > >
>> >> > > > > On Thursday, July 30th, 2026 at 3:43 PM, L. C. Hsieh <
>> [email protected]
>> >> > > >
>> >> > > > > wrote:
>> >> > > > >
>> >> > > > > Hi all,
>> >> > > > >
>> >> > > > > I'd like to start a discussion on a SPIP proposing an official
>> gateway
>> >> > > > > component for Spark Connect deployments.
>> >> > > > >
>> >> > > > > Problem: A Spark Connect client has to know the address of one
>> specific
>> >> > > > > server and stay with it for the lifetime of its session,
>> because
>> >> > > session
>> >> > > > > state lives in that server's driver. As soon as a deployment
>> runs more
>> >> > > than
>> >> > > > > one Spark Connect server — per-team servers, capacity,
>> redundancy —
>> >> > > there
>> >> > > > > is no built-in way to put a single address in front of the
>> fleet. A
>> >> > > generic
>> >> > > > > load balancer can't do this either: the routing key
>> (session_id) lives
>> >> > > > > inside the request protobuf, not in headers. So today every
>> team
>> >> > > writes its
>> >> > > > > own proxy glue, and everyone solves the same problem from
>> scratch.
>> >> > > > >
>> >> > > > > Proposal: The Spark Connect Gateway is a standalone,
>> wire-compatible
>> >> > > gRPC
>> >> > > > > proxy that sits in front of a pool of Spark Connect servers.
>> >> > > > >
>> >> > > > > Clients are untouched: the connection string points at the
>> gateway and
>> >> > > > > everything else is unchanged. The proposal adds nothing to and
>> changes
>> >> > > > > nothing in the Spark Connect protocol.
>> >> > > > >
>> >> > > > > SPIP doc:
>> >> > > > >
>> >> > >
>> https://docs.google.com/document/d/16RNnRJjSDBSlcgeBjVz_Mv-SkqvH33F9tIsR1Ii_P3c/edit?tab=t.0#heading=h.221xxblzqzu
>> >> > > > > JIRA: https://issues.apache.org/jira/browse/SPARK-58455
>> >> > > > >
>> >> > > > > Feedback is very welcome.
>> >> > > > >
>> >> > > > > Thanks,
>> >> > > > > Liang-Chi
>> >> > > > >
>> >> > > > >
>> >> > > > >
>> >> > > >
>> >> > >
>> >> > >
>> ---------------------------------------------------------------------
>> >> > > To unsubscribe e-mail: [email protected]
>> >> > >
>> >> > >
>> >> >
>> >> > --
>> >> > John Zhuge
>> >> >
>> >>
>> >> ---------------------------------------------------------------------
>> >> To unsubscribe e-mail: [email protected]
>> >>
>>
>> ---------------------------------------------------------------------
>> To unsubscribe e-mail: [email protected]
>>
>>
>
> --
> Adam Binford
>

Reply via email to