Agree with Jun On Wed, Jun 24, 2026 at 2:57 PM Jun Wang via dev <[email protected]> wrote:
> I currently lean towards Proposal 1. > > A FlattenOutPut option could work if it is resolved before type checking > and optimization. However, I would be cautious about allowing platform > implementations to override it, since the logical output type should remain > platform-independent. > > Trino does support the SQL ROW type, so the composite-type approach is > interesting. Still, relying on it would require checking connector, JDBC, > and cross-platform support, and probably extending Wayang’s type system. > For the current scope, a dedicated relational join seems to me the clearest > and most practical solution. > > Best, > > Jun > > On Wed, Jun 24, 2026 at 3:38 PM Juri Petersen <[email protected]> wrote: > > > Hi, > > I also think building a solution into our type system is the cleanest > way. > > A composite type that would allow a view as a Tuple2<?, ?> or a flat type > > could solve this already, if I am not missing something. > > Having multiple instances of the same logical operator just for the > > purpose of flattening seems counterintuitive. > > > > Best, > > Juri > > ________________________________ > > From: Yasser Idris <[email protected]> > > Sent: 24 June 2026 04:02 > > To: [email protected] <[email protected]> > > Subject: Re: [DISCUSS] output type of Join operator issue > > > > +1 for What Gabor said, I don't think we need a new join operator for > > relational dbs. I think this breaks the abstraction of the logical > > operator. I'm more inclined towards solution 2 but with a twist. The > > logical(wayang) join operator should have 2 modes > > (FlattenOutput=True/False), and it defaults to False. Every platform > > implementation of the join operator can override the default, but also > > every plan (ie. this can be exposed to the user). This is not specific to > > relational database. If this boolean is True the output is chained into > our > > brand new FlattenOperator which as Zoi said, has platform specific > > implementations. The Tuple 2 problem can be addressed with abstraction, i > > don't see it as an issue and actually this insertion of the > Flattenoperator > > happens before optimization since it's part of the Join definition. > > > > > > > > On Tue, Jun 23, 2026 at 12:37 PM Zoi Kaoudi <[email protected]> wrote: > > > > > Thank you Gabor. I was not aware about this composite type for > relational > > > databases. However, I doubt that newest engines such as Trino or Presto > > > support this but we could check that. > > > > > > I still also prefer to add a new operator to have things cleaner. Then > a > > > dataframe api can be built on top of that in an easier manner I think. > > > > > > Best > > > -- > > > Zoi > > > > > > > > > On 2026/06/23 08:55:38 Gábor Gévay wrote: > > > > Btw. many relational engines also have Record or Row or some kind of > > > > composite type, in which case it's actually possible to represent > > > > Tuple2<Record, Record> directly. So, a 3rd solution could be to rely > > > > on this. But the caveat is that I'm not sure how well this is > > > > supported across systems. Even though the SQL:1999 standard has it > > > > (calls it `ROW`), but unfortunately not every "standard" thing is > > > > implemented even by the mainstream systems. > > > > > > > > For example, in Postgres: > > > > ``` > > > > CREATE TYPE lhs AS (a int, b int); > > > > CREATE TYPE rhs AS (x int, y int); > > > > > > > > CREATE TABLE joined ( > > > > l lhs, > > > > r rhs > > > > ); > > > > > > > > INSERT INTO joined VALUES (ROW(1,2), ROW(3,4)); > > > > INSERT INTO joined VALUES (ROW(5,6), ROW(7,8)); > > > > > > > > SELECT * FROM joined; > > > > > > > > SELECT (l).*, (r).* FROM joined; > > > > ``` > > > > gives you: > > > > ``` > > > > l | r > > > > -------+------- > > > > (1,2) | (3,4) > > > > (5,6) | (7,8) > > > > (2 rows) > > > > > > > > a | b | x | y > > > > ---+---+---+--- > > > > 1 | 2 | 3 | 4 > > > > 5 | 6 | 7 | 8 > > > > (2 rows) > > > > ``` > > > > (The second SELECT shows how to flatten it.) > > > > > > > > Best, > > > > Gábor > > > > > > > > > > > > > > > > > > > > Alexander Alten <[email protected]> ezt írta (időpont: 2026. jún. > 23., > > > K, 10:13): > > > > > > > > > > Hi community, > > > > > > > > > > Good discussion - I’d support Kaustubh. It is an additional > operator > > > for a different case, and should be also defined as one. > > > > > > > > > > Best, > > > > > —Alex > > > > > > > > > > > On Jun 23, 2026, at 06:35, Kaustubh Beedkar <[email protected]> > > > wrote: > > > > > > > > > > > > Hi Zoi, > > > > > > > > > > > > This is exactly the issue I had to deal with when implementing > the > > > initial > > > > > > join operator in SQL api. > > > > > > > > > > > > The Tuple2<..> is a must for the dataflow case as it "encodes" > the > > > pairing > > > > > > of two inputs while preserving the sides. But of course this is > the > > > wrong > > > > > > algebra for the relational case. > > > > > > > > > > > > Prposal 1 seems like a better option IMO. > > > > > > > > > > > > Best > > > > > > Kaustubh > > > > > > > > > > > > On Mon, Jun 22, 2026 at 7:45 PM Zoi Kaoudi via dev < > > > [email protected]> > > > > > > wrote: > > > > > > > > > > > >> Dear all, > > > > > >> > > > > > >> this is a long email, please bare with me. > > > > > >> > > > > > >> We have a fundamental and recurring issue with the output type > of > > > the > > > > > >> current Join operator. The Join operator as is now takes as > input > > > two > > > > > >> generic datatypes (potentially different) and outputs a Tuple2 > of > > > these > > > > > >> input datatypes. See here: > > > > > >> > > > > > >> > > > > > >> > > > > > > https://urldefense.proofpoint.com/v2/url?u=https-3A__www.google.com_url-3Fq-3Dhttps-3A__github.com_apache_wayang_blob_cb4bb0fd0245240124174b562e4ad1abbc9cd0c0_wayang-2Dcommons_wayang-2Dbasic_src_main_java_org_apache_wayang_basic_operators_JoinOperator.java-2523L37-26source-3Dgmail-2Dimap-26ust-3D1782794153000000-26usg-3DAOvVaw3d4SXsHbG8-2DgRlC0sCSPNI&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=cj42uI7ljZf7BV9opO9YjAqil63_oZ3cCHcZB-Ad4cw&e= > > > > > >> > > > > > >> This is fine for a general dataflow type of engine. It also > > enables > > > > > >> different types of joins, like joining specific java object > > > (Employee with > > > > > >> Employee) with a custom join key. However, when we are handling > > > relational > > > > > >> data and have JDBC platforms for execution this creates > problems. > > > > > >> > > > > > >> *Problem 1* > > > > > >> When the inputs are Records, the join output is a Tuple2<Record, > > > Record> > > > > > >> rather than a flat Record. Any relational operator we chain > after > > > the join > > > > > >> therefore cannot consume the output directly; the user has to > > > insert a > > > > > >> flatten step (a MapOperator) to turn the Tuple2 into a single > > > Record. This > > > > > >> is very cumbersome of course. > > > > > >> > > > > > >> *Problem 2* > > > > > >> Even if we "hide" a flatten map operator step somehow and we > > insert > > > it > > > > > >> automatically, this step cannot run inside a database. So even > > when > > > the > > > > > >> whole pipeline could execute on a JDBC platform, the flatten > > > MapOperator > > > > > >> forces the optimizer to move data out of the database into the > JVM > > > just to > > > > > >> reshape it. This is ok as long as every in-database pipeline > ends > > by > > > > > >> streaming its results back to Java anyway. However, if we want > to > > > have > > > > > >> fully in-database pipelines, this becomes problematic: For > > example, > > > with an > > > > > >> in-database sink that writes the join result back into a table > > > (CREATE > > > > > >> TABLE AS SELECT) and expects a Record as input, connecting the > > join > > > to the > > > > > >> sink fails at plan construction because Tuple2 != Record, even > > > though the > > > > > >> SQL engine would naturally produce a flat row. > > > > > >> > > > > > >> *Proposed solution 1* > > > > > >> We introduce a new join operator for relational data that takes > > two > > > > > >> Records and outputs a single, flat Record. On JDBC this maps > > > directly to a > > > > > >> SQL join statement which already yields flat rows; on Java/Spark > > it > > > > > >> concatenates the two Records instead of wrapping them in a > Tuple2. > > > The > > > > > >> current Tuple2-based Join should stay for the general dataflow > > > case, and > > > > > >> the Record-based one is used for the relational case. The same > > idea > > > should > > > > > >> be extended to the other operators that share this wrapping > > > behaviour (e.g. > > > > > >> the cartesian product). > > > > > >> Pros: flat types end to end and in-database pipelines can stay > > > fully in > > > > > >> the database. > > > > > >> Cons: a second join operator to maintain and the relational/SQL > > > path has > > > > > >> to target it. > > > > > >> > > > > > >> *Proposed solution 2* > > > > > >> We keep the current Join and add a new dedicated Wayang flatten > > > operator > > > > > >> (eg FlattenOperator) that converts Tuple2<Record, Record> into a > > > flat > > > > > >> Record, with a platform-specific implementations: on JDBC it > > > essentially > > > > > >> adds an empty string (the SQL already produces flat rows, so it > > > contributes > > > > > >> nothing to the query), while on Java/Spark it performs the > actual > > > > > >> concatenation. Ideally the optimizer should insert it > > automatically > > > > > >> wherever the join will be executed. > > > > > >> Pros: reuses the existing, fully supported operators, with a > > single > > > > > >> mechanism. > > > > > >> Cons: we keep the Tuple2 in the logical plan and since the type > > > check > > > > > >> happens before optimization, the automatic insertion may need to > > > happen > > > > > >> somewhere else. > > > > > >> > > > > > >> I'd appreciate any thoughts on this and also any experiences you > > > had and > > > > > >> how you solved them or how you would like to have them solved. > > > > > >> Also feel free to propose some other solution. > > > > > >> > > > > > >> Best > > > > > >> -- > > > > > >> Zoi > > > > > >> > > > > > > > > > > > > > > > -- > > > > > *Scalytics* > > > > > The foundation for secure, scalable, and transparent AI. > > > > > > > > https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e= > > < > > > https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e= > > > < > > > https://urldefense.proofpoint.com/v2/url?u=http-3A__www.scalytics.io&d=DwIGaQ&c=slrrB7dE8n7gBJbeO0g-IQ&r=K-pGUp478bHBxJaHksONsg&m=qJpq2b8-Boxfz0iGjP52w8YwQ6PomnnhJepFaWpeogFjf9fQlhHGRSVNCv_iVJ19&s=Rd9RXVrx-cbXbHHUTUTmgfLBerwSxY5LSD6amtKGjvo&e= > > > > > > > > > > > > > -- Please consider the > > > > > environment before printing this email -- > > > > > > > > > > Disclaimer: > > > > > The content of this > > > > > message is confidential. If you have received it by mistake, please > > > inform > > > > > us by an email reply and then delete the message. It is forbidden > to > > > copy, > > > > > forward, or in any way reveal the contents of this message to > anyone. > > > The > > > > > integrity and security of this email cannot be guaranteed over the > > > > > Internet. Therefore, the sender will not be held liable for any > > damage > > > > > caused by the message. > > > > > > > > > >
