I currently lean towards Proposal 1. A FlattenOutPut option could work if it is resolved before type checking and optimization. However, I would be cautious about allowing platform implementations to override it, since the logical output type should remain platform-independent.
Trino does support the SQL ROW type, so the composite-type approach is interesting. Still, relying on it would require checking connector, JDBC, and cross-platform support, and probably extending Wayang’s type system. For the current scope, a dedicated relational join seems to me the clearest and most practical solution. Best, Jun On Wed, Jun 24, 2026 at 3:38 PM Juri Petersen <[email protected]> wrote: > Hi, > I also think building a solution into our type system is the cleanest way. > A composite type that would allow a view as a Tuple2<?, ?> or a flat type > could solve this already, if I am not missing something. > Having multiple instances of the same logical operator just for the > purpose of flattening seems counterintuitive. > > Best, > Juri > ________________________________ > From: Yasser Idris <[email protected]> > Sent: 24 June 2026 04:02 > To: [email protected] <[email protected]> > Subject: Re: [DISCUSS] output type of Join operator issue > > +1 for What Gabor said, I don't think we need a new join operator for > relational dbs. I think this breaks the abstraction of the logical > operator. I'm more inclined towards solution 2 but with a twist. The > logical(wayang) join operator should have 2 modes > (FlattenOutput=True/False), and it defaults to False. Every platform > implementation of the join operator can override the default, but also > every plan (ie. this can be exposed to the user). This is not specific to > relational database. If this boolean is True the output is chained into our > brand new FlattenOperator which as Zoi said, has platform specific > implementations. The Tuple 2 problem can be addressed with abstraction, i > don't see it as an issue and actually this insertion of the Flattenoperator > happens before optimization since it's part of the Join definition. > > > > On Tue, Jun 23, 2026 at 12:37 PM Zoi Kaoudi <[email protected]> wrote: > > > Thank you Gabor. I was not aware about this composite type for relational > > databases. However, I doubt that newest engines such as Trino or Presto > > support this but we could check that. > > > > I still also prefer to add a new operator to have things cleaner. Then a > > dataframe api can be built on top of that in an easier manner I think. > > > > Best > > -- > > Zoi > > > > > > On 2026/06/23 08:55:38 Gábor Gévay wrote: > > > Btw. many relational engines also have Record or Row or some kind of > > > composite type, in which case it's actually possible to represent > > > Tuple2<Record, Record> directly. So, a 3rd solution could be to rely > > > on this. But the caveat is that I'm not sure how well this is > > > supported across systems. Even though the SQL:1999 standard has it > > > (calls it `ROW`), but unfortunately not every "standard" thing is > > > implemented even by the mainstream systems. > > > > > > For example, in Postgres: > > > ``` > > > CREATE TYPE lhs AS (a int, b int); > > > CREATE TYPE rhs AS (x int, y int); > > > > > > CREATE TABLE joined ( > > > l lhs, > > > r rhs > > > ); > > > > > > INSERT INTO joined VALUES (ROW(1,2), ROW(3,4)); > > > INSERT INTO joined VALUES (ROW(5,6), ROW(7,8)); > > > > > > SELECT * FROM joined; > > > > > > SELECT (l).*, (r).* FROM joined; > > > ``` > > > gives you: > > > ``` > > > l | r > > > -------+------- > > > (1,2) | (3,4) > > > (5,6) | (7,8) > > > (2 rows) > > > > > > a | b | x | y > > > ---+---+---+--- > > > 1 | 2 | 3 | 4 > > > 5 | 6 | 7 | 8 > > > (2 rows) > > > ``` > > > (The second SELECT shows how to flatten it.) > > > > > > Best, > > > Gábor > > > > > > > > > > > > > > > Alexander Alten <[email protected]> ezt írta (időpont: 2026. jún. 23., > > K, 10:13): > > > > > > > > Hi community, > > > > > > > > Good discussion - I’d support Kaustubh. It is an additional operator > > for a different case, and should be also defined as one. > > > > > > > > Best, > > > > —Alex > > > > > > > > > On Jun 23, 2026, at 06:35, Kaustubh Beedkar <[email protected]> > > wrote: > > > > > > > > > > Hi Zoi, > > > > > > > > > > This is exactly the issue I had to deal with when implementing the > > initial > > > > > join operator in SQL api. > > > > > > > > > > The Tuple2<..> is a must for the dataflow case as it "encodes" the > > pairing > > > > > of two inputs while preserving the sides. But of course this is the > > wrong > > > > > algebra for the relational case. > > > > > > > > > > Prposal 1 seems like a better option IMO. > > > > > > > > > > Best > > > > > Kaustubh > > > > > > > > > > On Mon, Jun 22, 2026 at 7:45 PM Zoi Kaoudi via dev < > > [email protected]> > > > > > wrote: > > > > > > > > > >> Dear all, > > > > >> > > > > >> this is a long email, please bare with me. > > > > >> > > > > >> We have a fundamental and recurring issue with the output type of > > the > > > > >> current Join operator. The Join operator as is now takes as input > > two > > > > >> generic datatypes (potentially different) and outputs a Tuple2 of > > these > > > > >> input datatypes. See here: > > > > >> > > > > >> > > > > >> > > > https://urldefense.proofpoint.com/v2/url?u=https-3A__www.google.com_url-3Fq-3Dhttps-3A__github.com_apache_wayang_blob_cb4bb0fd0245240124174b562e4ad1abbc9cd0c0_wayang-2Dcommons_wayang-2Dbasic_src_main_java_org_apache_wayang_basic_operators_JoinOperator.java-2523L37-26source-3Dgmail-2Dimap-26ust-3D1782794153000000-26usg-3DAOvVaw3d4SXsHbG8-2DgRlC0sCSPNI&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=cj42uI7ljZf7BV9opO9YjAqil63_oZ3cCHcZB-Ad4cw&e= > > > > >> > > > > >> This is fine for a general dataflow type of engine. It also > enables > > > > >> different types of joins, like joining specific java object > > (Employee with > > > > >> Employee) with a custom join key. However, when we are handling > > relational > > > > >> data and have JDBC platforms for execution this creates problems. > > > > >> > > > > >> *Problem 1* > > > > >> When the inputs are Records, the join output is a Tuple2<Record, > > Record> > > > > >> rather than a flat Record. Any relational operator we chain after > > the join > > > > >> therefore cannot consume the output directly; the user has to > > insert a > > > > >> flatten step (a MapOperator) to turn the Tuple2 into a single > > Record. This > > > > >> is very cumbersome of course. > > > > >> > > > > >> *Problem 2* > > > > >> Even if we "hide" a flatten map operator step somehow and we > insert > > it > > > > >> automatically, this step cannot run inside a database. So even > when > > the > > > > >> whole pipeline could execute on a JDBC platform, the flatten > > MapOperator > > > > >> forces the optimizer to move data out of the database into the JVM > > just to > > > > >> reshape it. This is ok as long as every in-database pipeline ends > by > > > > >> streaming its results back to Java anyway. However, if we want to > > have > > > > >> fully in-database pipelines, this becomes problematic: For > example, > > with an > > > > >> in-database sink that writes the join result back into a table > > (CREATE > > > > >> TABLE AS SELECT) and expects a Record as input, connecting the > join > > to the > > > > >> sink fails at plan construction because Tuple2 != Record, even > > though the > > > > >> SQL engine would naturally produce a flat row. > > > > >> > > > > >> *Proposed solution 1* > > > > >> We introduce a new join operator for relational data that takes > two > > > > >> Records and outputs a single, flat Record. On JDBC this maps > > directly to a > > > > >> SQL join statement which already yields flat rows; on Java/Spark > it > > > > >> concatenates the two Records instead of wrapping them in a Tuple2. > > The > > > > >> current Tuple2-based Join should stay for the general dataflow > > case, and > > > > >> the Record-based one is used for the relational case. The same > idea > > should > > > > >> be extended to the other operators that share this wrapping > > behaviour (e.g. > > > > >> the cartesian product). > > > > >> Pros: flat types end to end and in-database pipelines can stay > > fully in > > > > >> the database. > > > > >> Cons: a second join operator to maintain and the relational/SQL > > path has > > > > >> to target it. > > > > >> > > > > >> *Proposed solution 2* > > > > >> We keep the current Join and add a new dedicated Wayang flatten > > operator > > > > >> (eg FlattenOperator) that converts Tuple2<Record, Record> into a > > flat > > > > >> Record, with a platform-specific implementations: on JDBC it > > essentially > > > > >> adds an empty string (the SQL already produces flat rows, so it > > contributes > > > > >> nothing to the query), while on Java/Spark it performs the actual > > > > >> concatenation. Ideally the optimizer should insert it > automatically > > > > >> wherever the join will be executed. > > > > >> Pros: reuses the existing, fully supported operators, with a > single > > > > >> mechanism. > > > > >> Cons: we keep the Tuple2 in the logical plan and since the type > > check > > > > >> happens before optimization, the automatic insertion may need to > > happen > > > > >> somewhere else. > > > > >> > > > > >> I'd appreciate any thoughts on this and also any experiences you > > had and > > > > >> how you solved them or how you would like to have them solved. > > > > >> Also feel free to propose some other solution. > > > > >> > > > > >> Best > > > > >> -- > > > > >> Zoi > > > > >> > > > > > > > > > > > > -- > > > > *Scalytics* > > > > The foundation for secure, scalable, and transparent AI. > > > > > https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e= > < > https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e= > > < > https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e= > > > > > > > > > > -- Please consider the > > > > environment before printing this email -- > > > > > > > > Disclaimer: > > > > The content of this > > > > message is confidential. If you have received it by mistake, please > > inform > > > > us by an email reply and then delete the message. It is forbidden to > > copy, > > > > forward, or in any way reveal the contents of this message to anyone. > > The > > > > integrity and security of this email cannot be guaranteed over the > > > > Internet. Therefore, the sender will not be held liable for any > damage > > > > caused by the message. > > > > > >
