Pranshu-S commented on PR #16710:
URL: https://github.com/apache/lucene/pull/16710#issuecomment-5935737988

   I see, so with `index-time-filter` we are validating the gains with dedup 
compared to current setup with the expectation that recall should be same with 
improvements in index_size?
   
   Looks to be pretty much within the expectations -
   (Note - filter selectivity = 0.5)
   
   ```
   
==================================================================================
   COMPARISON  (candidate=lucene_candidate+dedup vs 
baseline=lucene_baseline+non-dedup;
                delta/pct = candidate - baseline)
   
==================================================================================
   metric             baseline/non-dedup(n=1)  candidate/dedup(n=1)  delta     
pct
   -----------------  -----------------------  --------------------  --------  
------
   recall             0.976                    0.976                 +0.000    
+0.0%
   latency(ms)        0.566                    0.592                 +0.026    
+4.6%
   netCPU             0.559                    0.586                 +0.027    
+4.8%
   avgCpuCount        0.988                    0.990                 +0.002    
+0.2%
   nDoc               10000                    10000
   searchType         KNN                      KNN
   topK               10                       10
   fanout             100                      100
   resultSimilarity   N/A                      N/A
   decay              N/A                      N/A
   resultCount        10.000                   10.000
   maxConn            32                       32
   beamWidth          200                      200
   quantized          7 bits                   7 bits
   visited            2122.000                 2115.000              -7.000    
-0.3%
   index(s)           3.300                    3.090                 -0.210    
-6.4%
   index_docs/s       3028.470                 3234.150              +205.680  
+6.8%
   merge(s)           0.000                    0.000                 +0.000
   force_merge(s)     11.160                   10.300                -0.860    
-7.7%
   num_segments       1.000                    1.000                 +0.000    
+0.0%
   index_size(MB)     74.230                   49.930                -24.300   
-32.7%
   filterStrategy     index-time-filter        index-time-filter
   filterSelectivity  0.50                     0.50
   overSample         1.000                    1.000
   vec_disk(MB)       48.981                   48.981                +0.000    
+0.0%
   vec_RAM(MB)        9.918                    9.918                 +0.000    
+0.0%
   bp-reorder         false                    false
   indexType          HNSW                     HNSW
   rerank             no                       no
   
   ```


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to