[ 
https://issues.apache.org/jira/browse/PHOENIX-7998?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Andrew Kyle Purtell updated PHOENIX-7998:
-----------------------------------------
    Description: 
Vector similarity search uses mathematical embeddings to identify records by 
semantic meaning and conceptual proximity rather than exact keyword matches. In 
a hybrid document database like Apache Phoenix, native vector search allows 
applications to execute unified queries that combine AI-driven semantic 
retrieval with relational filters and nested BSON document predicates. This 
eliminates the operational overhead, data replication, and synchronization lag 
of maintaining external vector databases while leveraging Phoenix's existing 
scale-out query parallelism.

This design integrates native vector similarity search into Apache Phoenix 
across relational columns and BSON documents using fixed-dimension vector 
types, typed path extraction expressions, and standard SQL distance functions. 
Exact search pushes distance evaluation and bounded candidate heaps down to 
RegionServers for distributed top-N execution, while approximate retrieval maps 
Inverted File centroid posting lists into contiguous global index ranges 
traversed via multi-range skip scans. Server-side coprocessors maintain 
centroid assignments and verified-write consistency across relational mutations 
and atomic BSON updates, background compactions observe centroid drift to 
trigger asynchronous rebuilds, and the optimizer balances exact evaluation, 
covered index scanning, and adaptive probing within an extensible architecture.

Design Document: 
https://docs.google.com/document/d/139E87IC2b6UG5T0UaxOK64qAjPt2u-ePB7DirWDm3v0/edit?tab=t.0

Design Document as gist, may be more recent: 
https://gist.github.com/apurtell/3af244c582d94ba83c61d5eca2c8b8c5


  was:
Vector similarity search uses mathematical embeddings to identify records by 
semantic meaning and conceptual proximity rather than exact keyword matches. In 
a hybrid document database like Apache Phoenix, native vector search allows 
applications to execute unified queries that combine AI-driven semantic 
retrieval with relational filters and nested BSON document predicates. This 
eliminates the operational overhead, data replication, and synchronization lag 
of maintaining external vector databases while leveraging Phoenix's existing 
scale-out query parallelism.

This design integrates native vector similarity search into Apache Phoenix 
across relational columns and BSON documents using fixed-dimension vector 
types, typed path extraction expressions, and standard SQL distance functions. 
Exact search pushes distance evaluation and bounded candidate heaps down to 
RegionServers for distributed top-N execution, while approximate retrieval maps 
Inverted File centroid posting lists into contiguous global index ranges 
traversed via multi-range skip scans. Server-side coprocessors maintain 
centroid assignments and verified-write consistency across relational mutations 
and atomic BSON updates, background compactions observe centroid drift to 
trigger asynchronous rebuilds, and the optimizer balances exact evaluation, 
covered index scanning, and adaptive probing within an extensible architecture.

Design Document: 
https://docs.google.com/document/d/139E87IC2b6UG5T0UaxOK64qAjPt2u-ePB7DirWDm3v0/edit?tab=t.0


> Vector Indexes
> --------------
>
>                 Key: PHOENIX-7998
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-7998
>             Project: Phoenix
>          Issue Type: New Feature
>          Components: core
>            Reporter: Andrew Kyle Purtell
>            Assignee: Andrew Kyle Purtell
>            Priority: Major
>
> Vector similarity search uses mathematical embeddings to identify records by 
> semantic meaning and conceptual proximity rather than exact keyword matches. 
> In a hybrid document database like Apache Phoenix, native vector search 
> allows applications to execute unified queries that combine AI-driven 
> semantic retrieval with relational filters and nested BSON document 
> predicates. This eliminates the operational overhead, data replication, and 
> synchronization lag of maintaining external vector databases while leveraging 
> Phoenix's existing scale-out query parallelism.
> This design integrates native vector similarity search into Apache Phoenix 
> across relational columns and BSON documents using fixed-dimension vector 
> types, typed path extraction expressions, and standard SQL distance 
> functions. Exact search pushes distance evaluation and bounded candidate 
> heaps down to RegionServers for distributed top-N execution, while 
> approximate retrieval maps Inverted File centroid posting lists into 
> contiguous global index ranges traversed via multi-range skip scans. 
> Server-side coprocessors maintain centroid assignments and verified-write 
> consistency across relational mutations and atomic BSON updates, background 
> compactions observe centroid drift to trigger asynchronous rebuilds, and the 
> optimizer balances exact evaluation, covered index scanning, and adaptive 
> probing within an extensible architecture.
> Design Document: 
> https://docs.google.com/document/d/139E87IC2b6UG5T0UaxOK64qAjPt2u-ePB7DirWDm3v0/edit?tab=t.0
> Design Document as gist, may be more recent: 
> https://gist.github.com/apurtell/3af244c582d94ba83c61d5eca2c8b8c5



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to