+1 to adding this. It's a really awesome idea. It really fits well
into the approximate search support that AsterixDB has featured and
integrated deeply, ever since the first open source release
(pre-apache). Being able to express something that is usually a giant
imperative pipeline as a simple declarative expression is very unique.

On Mon, Aug 17, 2026 at 11:35 AM Hongyu Shi <[email protected]> wrote:
>
> Hi all, Initiating discussion to add cluster-by clause in AsterixDB.
> Feature: Cluster-By Clause
>
> While GROUP BY is the canonical clause for analytic queries to further
> understand the trend and nature of the underlying data, our proposed
> Cluster-By clause is for users to find the patterns and trends of vector
> embeddings by putting similar vector embeddings into some number of
> clusters.
>
> Clustering (CLUSTER BY) functionality for vector embeddings is similar to
> grouping (GROUP BY) for structured fields with the difference that
> clustering for vector embeddings is a fuzzy and similarity-based clustering
> while grouping of structured fields is exact value grouping. Due to the
> similarities between the two approaches, clustering functionality is
> proposed here to be added to SQL++ syntax using CLUSTER BY and CLUSTER AS,
> following the design and syntax of existing GROUP BY and GROUP AS clauses
> in SQL++.
> APE:
> https://cwiki.apache.org/confluence/spaces/ASTERIXDB/pages/406621607/APE+33+CLUSTER+BY
> Thanks, Hongyu Shi

Reply via email to