airborne12 opened a new pull request, #68758:
URL: https://github.com/apache/doris/pull/68758
> Performance follow-up remains open: V2 exact-term SQL is +7.67% and V2
SEARCH phrase SQL is +7.65% against the frozen baseline. Details and pending
validation are listed below.
### What problem does this PR solve?
Issue Number: None
Related PR: #67180, #67538, #67918
MATCH and SEARCH maintained separate Boolean, phrase, term-expansion and
scoring implementations for CLucene/V2 and SNII. Common query optimizations
required changes in multiple executors, and cross-field queries could combine
document IDs from incompatible physical partitions.
This change lowers both interfaces into common logical nodes and executes
them through shared Boolean, phrase and expansion code. Format adapters supply
dictionaries, postings, positions, norms and I/O. Candidates, NULL bitmaps and
collection use global document IDs; CLucene partition traversal stays inside
its postings adapter. Obsolete executors and unused compiler/planning
interfaces are removed.
Unscored conjunctions order ordinary clauses and propagate TRUE and UNKNOWN
candidates to deferred expansions. Exact phrases use shared lazy verification,
while repeated terms, sloppy phrases and scored frequencies retain the
positions and native norms they need. The common text reader owns
cache/coalescing, count-only lookup and candidate-consumption decisions. Cache
keys own one encoded buffer; native document advancement avoids redundant work
and status wrappers.
Unscored root expansions collect directly into caller output through the
same collector used by SEARCH expansion weights. One source operation collects
terms: V2 uses existing cursors and SNII its existing bulk document decoder.
This removes unnecessary Query/Weight/context/scorer and temporary-bitmap
state. Candidate intersections remain shared. Borrowed dictionary entries
remain source-owned through synchronous reads; the unused vector-of-vectors
decoder is removed. A real-reader regression verifies that direct collection
preserves the existing V2 error conversion.
Sparse scored-OR merging is bounded. Physical batch-read buffers are
retained through owned views, removing scatter copies without changing read
ranges, coalescing or scheduling. These views retain existing coalescing gaps
until a wave ends; lower peak memory is not universal. Writer algorithms and
persisted formats are unchanged. There are no benchmark-specific branches,
tuned thresholds, compiler/allocator changes or new persistent caches.
Future unscored Boolean ordering, candidate propagation, expansion and
phrase verification can be optimized once for V2 and SNII. Scored conjunctions
still construct scorers eagerly and must cover batched ListedTerms and
streaming intersections. Combining independent SQL MATCH predicates needs
planner/SegmentIterator integration with NULL handling; physical codec/I/O work
remains format-specific.
### Release note
- SEARCH REGEXP on SNII matches whole terms, consistently with V2.
- SEARCH PREFIX and WILDCARD normalize terms without tokenizing patterns.
- Default lucene-mode SEARCH excludes a NULL field from the positive clause,
allowing NOT and MUST_NOT to retain that row. Standard mode preserves SQL
three-valued logic.
- SNII SEARCH supports TERM minimum_should_match and multiple terms at one
phrase position. SNII MATCH_PHRASE supports stacked analyzer tokens.
- SNII full-text indexes support MATCH_PHRASE_EDGE; CLucene keyword indexes
evaluate it by function. V2 keyword MATCH_PHRASE recognizes a trailing slop
suffix.
- CLucene text indexes support count-only dictionary lookup of one exact
term and concurrent identical-query coalescing.
- SNII default-OR SEARCH terms publish BM25 scores instead of the previous
constant-bitset fallback. Scored cross-field OR retains a row matched by one
field when another field is NULL.
- Cross-field queries preserve document IDs across different physical
partitions. Materialized phrase postings preserve native norms for BM25, and
CLucene expansion limits apply across the field.
- SNII full-frame position profiles preserve the logical selection. Seven
counters used only by deleted SNII executors are removed.
The existing missing-term NULL asymmetry remains: a missing SEARCH term can
leave NULL rows UNKNOWN on CLucene and FALSE on SNII.
### Check List (For Author)
- Test
- [ ] Regression test: the final 48 affected ASAN cluster suites and 16
global-document-domain oracles remain pending.
- [x] Unit Test: 1677 ASAN cases pass before the final error-boundary
repair; the affected 251-case scope passes after it. These scopes overlap and
are not summed.
- [x] Manual test: complete native performance matrix, production SQL
comparisons, result oracles and artifact checks described below.
**Performance acceptance is not complete. Two V2 production SQL cases exceed
the agreed 5% median tolerance and will be optimized in follow-up commits on
this PR.** Opening this PR does not mean these regressions are accepted for
merge.
| Current production SQL case | Candidate/master median ratio | Change | 95%
interval |
| --- | ---: | ---: | --- |
| V2 exact term (`body MATCH_ANY '424242'`) | 1.07674 | +7.67% | [0.97460,
1.22613] |
| V2 SEARCH phrase (`search('body:"retry attempt"')`) | 1.07653 | +7.65% |
[1.03891, 1.10973] |
| SNII phrase with an ID range | 1.01760 | +1.76% | [0.96388, 1.05750] |
These ratios measure the index-filter profile time, using the median of
paired candidate/master ratios. The other 16 of the original 18 primary SQL
gates pass; all 18 client-wall median gates pass. All 4608 measured result
oracles, 16 process identities and 16 persisted replica files pass validation.
Host CPU during these query windows is approximately 6–12%, with no swap,
reclaim, allocation stalls, major faults or OOM. The V2 failures are retained
and are not dismissed as host-load noise.
All 565 current-source native cases pass the same median-ratio threshold,
retaining all 149248 paired samples and 128 fresh process replicas. The maximum
median ratio is 1.04682. Both candidate global expansion-limit checks and exact
persisted-input checks pass. Confidence intervals are reported separately and
some extend above 1.05; the acceptance rule applies to medians.
The four additional production expansion SQL filter ratios are V2 full
0.54924, V2 range 0.58035, SNII full 0.98434 and SNII range 1.03552. All four
filter and client-wall median gates pass, together with 1024 query oracles, 16
process identities and 16 exact replica files.
The SQL protocol uses two local backends with byte-identical replicas, fresh
processes for eight pairs, alternating binary placement, per-query AB/BA
ordering, four warmups and 16 samples per case/role. Both backends use CPUs
20–23 and default worker counts. Profiles verify the selected execution host. A
same-binary control on the previous candidate passes this protocol's median
gates. Every measured sample is retained without load correction or optional
extension. Passive host observations cover the measured windows.
The performance baseline is Apache Doris commit
`7ac0eb2fdb426595f810e94ab77b7aa2eece12d5`; the published candidate is
`a8ed35ad7822954ff78907796b8f862a4221eb51`. The current production executable
has SHA-256 `6929a0494128d19ac3c03e72dfefee564ef977e0aa51f2b78ce1c5a012638650`;
the native benchmark has SHA-256
`38aadb17a665f1d364a65ce57570883c2dd641f1f5590166e4d7190aa10bd66d`. The
published BE source matches the measured source; the final support-script edit
only simplifies a comment.
Earlier construction validation passes all 12 metrics over 320 fresh
processes. All 640 persisted-file sizes match; SNII bytes match exactly and
CLucene differs only in its existing segment-version field, with container gaps
checked. Subsequent fixes change only query/read paths. Construction includes
fixture setup and does not isolate writer throughput.
Master has existing correctness limitations for SNII default-OR TERM scores,
scored cross-field OR with a NULL secondary field and globally capped prefix
over multiple CLucene partitions. Comparable native timings use equivalent
single-token AND clauses or a shared NULL-free secondary corpus where
necessary. Original candidate semantics retain correctness tests;
multi-partition capped prefix is checked separately.
All 14 changed C++ files have passing static coverage. Format16, header
hygiene, standard RELEASE and shipped ELF GLIBC_2.17 compatibility pass. Four
external-corpus ASAN cases are skipped and 15 existing cases disabled in the
initial run; the final affected run has no skips or disabled cases. Full
committed format and English checks pass against target
`75d1f03a7cc0654cc8e48e5318ae2631e96cd980`: clang-format 16 covers the complete
merged BE/Cloud source and test scope, and the English check covers all 263
changed source files. FE/generated-schema paths are unchanged, so the CI
Checkstyle trigger does not apply. The reviewed language exceptions cover 17
individual Unicode input literals, with none for comments or diagnostics.
- Behavior changed:
- [ ] No.
- [x] Yes. See the release note.
- Does this need documentation?
- [x] No.
- [ ] Yes.
### Check List (For Reviewer who merge this PR)
- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Add branch pick label
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]