I currently lean towards Proposal 1.

A FlattenOutPut option could work if it is resolved before type checking
and optimization. However, I would be cautious about allowing platform
implementations to override it, since the logical output type should remain
platform-independent.

Trino does support the SQL ROW type, so the composite-type approach is
interesting. Still, relying on it would require checking connector, JDBC,
and cross-platform support, and probably extending Wayang’s type system.
For the current scope, a dedicated relational join seems to me the clearest
and most practical solution.

Best,

Jun

On Wed, Jun 24, 2026 at 3:38 PM Juri Petersen <[email protected]> wrote:

> Hi,
> I also think building a solution into our type system is the cleanest way.
> A composite type that would allow a view as a Tuple2<?, ?> or a flat type
> could solve this already, if I am not missing something.
> Having multiple instances of the same logical operator just for the
> purpose of flattening seems counterintuitive.
>
> Best,
> Juri
> ________________________________
> From: Yasser Idris <[email protected]>
> Sent: 24 June 2026 04:02
> To: [email protected] <[email protected]>
> Subject: Re: [DISCUSS] output type of Join operator issue
>
> +1 for What Gabor said, I don't think we need a new join operator for
> relational dbs. I think this breaks the abstraction of the logical
> operator. I'm more inclined towards solution 2 but with a twist. The
> logical(wayang) join operator should have 2 modes
> (FlattenOutput=True/False), and it defaults to False. Every platform
> implementation of the join operator can override the default, but also
> every plan (ie. this can be exposed to the user). This is not specific to
> relational database. If this boolean is True the output is chained into our
> brand new FlattenOperator which as Zoi said, has platform specific
> implementations. The Tuple 2 problem can be addressed with abstraction, i
> don't see it as an issue and actually this insertion of the Flattenoperator
> happens before optimization since it's part of the Join definition.
>
>
>
> On Tue, Jun 23, 2026 at 12:37 PM Zoi Kaoudi <[email protected]> wrote:
>
> > Thank you Gabor. I was not aware about this composite type for relational
> > databases. However, I doubt that newest engines such as Trino or Presto
> > support this but we could check that.
> >
> > I still also prefer to add a new operator to have things cleaner. Then a
> > dataframe api can be built on top of that in an easier manner I think.
> >
> > Best
> > --
> > Zoi
> >
> >
> > On 2026/06/23 08:55:38 Gábor Gévay wrote:
> > > Btw. many relational engines also have Record or Row or some kind of
> > > composite type, in which case it's actually possible to represent
> > > Tuple2<Record, Record> directly. So, a 3rd solution could be to rely
> > > on this. But the caveat is that I'm not sure how well this is
> > > supported across systems. Even though the SQL:1999 standard has it
> > > (calls it `ROW`), but unfortunately not every "standard" thing is
> > > implemented even by the mainstream systems.
> > >
> > > For example, in Postgres:
> > > ```
> > > CREATE TYPE lhs AS (a int, b int);
> > > CREATE TYPE rhs AS (x int, y int);
> > >
> > > CREATE TABLE joined (
> > >   l lhs,
> > >   r rhs
> > > );
> > >
> > > INSERT INTO joined VALUES (ROW(1,2), ROW(3,4));
> > > INSERT INTO joined VALUES (ROW(5,6), ROW(7,8));
> > >
> > > SELECT * FROM joined;
> > >
> > > SELECT (l).*, (r).* FROM joined;
> > > ```
> > > gives you:
> > > ```
> > >    l   |   r
> > > -------+-------
> > >  (1,2) | (3,4)
> > >  (5,6) | (7,8)
> > > (2 rows)
> > >
> > >  a | b | x | y
> > > ---+---+---+---
> > >  1 | 2 | 3 | 4
> > >  5 | 6 | 7 | 8
> > > (2 rows)
> > > ```
> > > (The second SELECT shows how to flatten it.)
> > >
> > > Best,
> > > Gábor
> > >
> > >
> > >
> > >
> > > Alexander Alten <[email protected]> ezt írta (időpont: 2026. jún. 23.,
> > K, 10:13):
> > > >
> > > > Hi community,
> > > >
> > > > Good discussion - I’d support Kaustubh. It is an additional operator
> > for a different case, and should be also defined as one.
> > > >
> > > > Best,
> > > > —Alex
> > > >
> > > > > On Jun 23, 2026, at 06:35, Kaustubh Beedkar <[email protected]>
> > wrote:
> > > > >
> > > > > Hi Zoi,
> > > > >
> > > > > This is exactly the issue I had to deal with when implementing the
> > initial
> > > > > join operator in SQL api.
> > > > >
> > > > > The Tuple2<..> is a must for the dataflow case as it "encodes" the
> > pairing
> > > > > of two inputs while preserving the sides. But of course this is the
> > wrong
> > > > > algebra for the relational case.
> > > > >
> > > > > Prposal 1 seems like a better option IMO.
> > > > >
> > > > > Best
> > > > > Kaustubh
> > > > >
> > > > > On Mon, Jun 22, 2026 at 7:45 PM Zoi Kaoudi via dev <
> > [email protected]>
> > > > > wrote:
> > > > >
> > > > >> Dear all,
> > > > >>
> > > > >> this is a long email, please bare with me.
> > > > >>
> > > > >> We have a fundamental and recurring issue with the output type of
> > the
> > > > >> current Join operator. The Join operator as is now takes as input
> > two
> > > > >> generic datatypes (potentially different) and outputs a Tuple2 of
> > these
> > > > >> input datatypes. See here:
> > > > >>
> > > > >>
> > > > >>
> >
> https://urldefense.proofpoint.com/v2/url?u=https-3A__www.google.com_url-3Fq-3Dhttps-3A__github.com_apache_wayang_blob_cb4bb0fd0245240124174b562e4ad1abbc9cd0c0_wayang-2Dcommons_wayang-2Dbasic_src_main_java_org_apache_wayang_basic_operators_JoinOperator.java-2523L37-26source-3Dgmail-2Dimap-26ust-3D1782794153000000-26usg-3DAOvVaw3d4SXsHbG8-2DgRlC0sCSPNI&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=cj42uI7ljZf7BV9opO9YjAqil63_oZ3cCHcZB-Ad4cw&e=
> > > > >>
> > > > >> This is fine for a general dataflow type of engine. It also
> enables
> > > > >> different types of joins, like joining specific java object
> > (Employee with
> > > > >> Employee) with a custom join key. However, when we are handling
> > relational
> > > > >> data and have JDBC platforms for execution this creates problems.
> > > > >>
> > > > >> *Problem 1*
> > > > >> When the inputs are Records, the join output is a Tuple2<Record,
> > Record>
> > > > >> rather than a flat Record. Any relational operator we chain after
> > the join
> > > > >> therefore cannot consume the output directly; the user has to
> > insert a
> > > > >> flatten step (a MapOperator) to turn the Tuple2 into a single
> > Record. This
> > > > >> is very cumbersome of course.
> > > > >>
> > > > >> *Problem 2*
> > > > >> Even if we "hide" a flatten map operator step somehow and we
> insert
> > it
> > > > >> automatically, this step cannot run inside a database. So even
> when
> > the
> > > > >> whole pipeline could execute on a JDBC platform, the flatten
> > MapOperator
> > > > >> forces the optimizer to move data out of the database into the JVM
> > just to
> > > > >> reshape it. This is ok as long as every in-database pipeline ends
> by
> > > > >> streaming its results back to Java anyway. However, if we want to
> > have
> > > > >> fully in-database pipelines, this becomes problematic: For
> example,
> > with an
> > > > >> in-database sink that writes the join result back into a table
> > (CREATE
> > > > >> TABLE AS SELECT) and expects a Record as input, connecting the
> join
> > to the
> > > > >> sink fails at plan construction because Tuple2 != Record, even
> > though the
> > > > >> SQL engine would naturally produce a flat row.
> > > > >>
> > > > >> *Proposed solution 1*
> > > > >> We introduce a new join operator for relational data that takes
> two
> > > > >> Records and outputs a single, flat Record. On JDBC this maps
> > directly to a
> > > > >> SQL join statement which already yields flat rows; on Java/Spark
> it
> > > > >> concatenates the two Records instead of wrapping them in a Tuple2.
> > The
> > > > >> current Tuple2-based Join should stay for the general dataflow
> > case, and
> > > > >> the Record-based one is used for the relational case. The same
> idea
> > should
> > > > >> be extended to the other operators that share this wrapping
> > behaviour (e.g.
> > > > >> the cartesian product).
> > > > >> Pros: flat types end to end and in-database pipelines can stay
> > fully in
> > > > >> the database.
> > > > >> Cons: a second join operator to maintain and the relational/SQL
> > path has
> > > > >> to target it.
> > > > >>
> > > > >> *Proposed solution 2*
> > > > >> We keep the current Join and add a new dedicated Wayang flatten
> > operator
> > > > >> (eg FlattenOperator) that converts Tuple2<Record, Record> into a
> > flat
> > > > >> Record, with a platform-specific implementations: on JDBC it
> > essentially
> > > > >> adds an empty string (the SQL already produces flat rows, so it
> > contributes
> > > > >> nothing to the query), while on Java/Spark it performs the actual
> > > > >> concatenation. Ideally the optimizer should insert it
> automatically
> > > > >> wherever the join will be executed.
> > > > >> Pros: reuses the existing, fully supported operators, with a
> single
> > > > >> mechanism.
> > > > >> Cons: we keep the Tuple2 in the logical plan and since the type
> > check
> > > > >> happens before optimization, the automatic insertion may need to
> > happen
> > > > >> somewhere else.
> > > > >>
> > > > >> I'd appreciate any thoughts on this and also any experiences you
> > had and
> > > > >> how you solved them or how you would like to have them solved.
> > > > >> Also feel free to propose some other solution.
> > > > >>
> > > > >> Best
> > > > >> --
> > > > >> Zoi
> > > > >>
> > > >
> > > >
> > > > --
> > > > *Scalytics*
> > > > The foundation for secure, scalable, and transparent AI.
> > > >
> https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e=
> <
> https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e=
> > <
> https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e=
> >
> > > >
> > > > --  Please consider the
> > > > environment before printing this email --
> > > >
> > > > Disclaimer:
> > > > The content of this
> > > > message is confidential. If you have received it by mistake, please
> > inform
> > > > us by an email reply and then delete the message. It is forbidden to
> > copy,
> > > > forward, or in any way reveal the contents of this message to anyone.
> > The
> > > > integrity and security of this email cannot be guaranteed over the
> > > > Internet. Therefore, the sender will not be held liable for any
> damage
> > > > caused by the message.
> > >
> >
>

Reply via email to