This is an automated email from the ASF dual-hosted git repository.
tisonkun pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/datasketches-rust.git
The following commit(s) were added to refs/heads/main by this push:
new 8f58d09 test: organize benchmarks and integration tests as workspace
crates (#230)
8f58d09 is described below
commit 8f58d09789bc5f666e8fcdf2f052d9ae9b7b0565
Author: tison <[email protected]>
AuthorDate: Wed Aug 26 18:32:25 2026 +0800
test: organize benchmarks and integration tests as workspace crates (#230)
* test: move integration tests into workspace crate
* bench: organize benchmarks by sketch workload
---
CONTRIBUTING.md | 33 +-
Cargo.lock | 20 +-
Cargo.toml | 2 +-
benchmarks/Cargo.toml | 35 ++
.../bloom_test/main.rs => benchmarks/cpc/mod.rs | 2 +-
benchmarks/cpc/serde.rs | 56 +++
.../tests/bloom_test => benchmarks}/main.rs | 12 +-
benchmarks/tdigest/compress.rs | 61 +++
benchmarks/tdigest/merge.rs | 102 +++++
.../hll_test/main.rs => benchmarks/tdigest/mod.rs | 6 +-
benchmarks/tdigest/query.rs | 58 +++
benchmarks/tdigest/serde.rs | 138 +++++++
benchmarks/tdigest/support.rs | 114 ++++++
benchmarks/tdigest/update.rs | 115 ++++++
datasketches/Cargo.toml | 53 ---
datasketches/benches/tdigest.rs | 456 ---------------------
tests-integration/Cargo.toml | 48 +++
.../main.rs => tests-integration/src/lib.rs | 2 +-
.../tests/bloom_test/main.rs | 0
.../tests/bloom_test/sketch.rs | 0
.../tests/countmin_test/main.rs | 0
.../tests/countmin_test/sketch.rs | 0
.../tests/cpc_test/deserialize.rs | 0
.../tests/cpc_test/main.rs | 0
.../tests/cpc_test/union.rs | 0
.../tests/cpc_test/update.rs | 0
.../tests/cpc_test/wrapper.rs | 0
.../tests/frequencies_test/main.rs | 0
.../tests/frequencies_test/update.rs | 0
.../tests/hll_test/main.rs | 0
.../tests/hll_test/union.rs | 0
.../tests/hll_test/update.rs | 0
.../tests/req_test/accuracy.rs | 0
.../tests/req_test/bounds.rs | 0
.../tests/req_test/core.rs | 0
.../tests/req_test/main.rs | 0
.../tests/req_test/merge.rs | 0
.../tests/req_test/property.rs | 0
.../tests/req_test/query.rs | 0
.../tests/req_test/sorted_view_api.rs | 0
.../tests/req_test/structure.rs | 0
.../tests/req_test/union.rs | 0
.../tests/serde_tests.rs | 9 -
.../tests/serde_tests/.gitignore | 0
.../tests/serde_tests/bloom.rs | 0
.../tests/serde_tests/countmin.rs | 0
.../tests/serde_tests/cpc.rs | 0
.../tests/serde_tests/frequencies.rs | 0
.../tests/serde_tests/hll.rs | 0
.../tdigest_ref_k100_n10000_double.sk | Bin
.../tdigest_ref_k100_n10000_float.sk | Bin
.../tests/serde_tests/req.rs | 0
.../tests/serde_tests/tdigest.rs | 0
.../tests/serde_tests/theta.rs | 0
.../tests/serde_tests/tuple.rs | 0
.../tests/tdigest_test/main.rs | 0
.../tests/tdigest_test/sketch.rs | 0
.../tests/theta_test/a_not_b.rs | 0
.../tests/theta_test/intersection.rs | 0
.../tests/theta_test/jaccard_similarity.rs | 0
.../tests/theta_test/main.rs | 0
.../tests/theta_test/sketch.rs | 0
.../tests/theta_test/union.rs | 0
.../tests/tuple_test/a_not_b.rs | 0
.../tests/tuple_test/intersection.rs | 0
.../tests/tuple_test/jaccard_similarity.rs | 0
.../tests/tuple_test/main.rs | 0
.../tests/tuple_test/sketch.rs | 0
.../tests/tuple_test/union.rs | 0
xtask/src/main.rs | 4 +-
70 files changed, 782 insertions(+), 544 deletions(-)
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
index dcbe46f..6634520 100644
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -66,6 +66,14 @@ Benchmark:
cargo x bench
```
+Benchmarks live in the standalone `benchmarks` crate. Files are grouped first
by sketch and then by workload, for example `benchmarks/cpc/serde.rs` and
`benchmarks/tdigest/update.rs`. Keep distinct workloads in separate modules so
Divan reports stable names such as `cpc::serde::serialize`.
+
+To run only one sketch or workload, pass a Divan filter directly:
+
+```shell
+cargo bench --package benchmarks --bench benchmarks -- cpc::serde
+```
+
## Public API documentation
- Describe types with noun phrases and API behavior with third-person
present-tense verbs such as `Creates`, `Updates`, and `Returns`.
@@ -82,28 +90,19 @@ cargo x bench
## Integration test layout
-Integration tests for the `datasketches` crate live under `datasketches/tests`
and use two entry-point patterns.
+End-to-end tests live in the standalone `tests-integration` crate, which
depends on `datasketches` with every sketch feature enabled. Unit tests that
require private implementation access remain next to the library code.
### Sketch behavior tests
-Non-serialization tests are grouped into one integration-test target per
sketch. The target entry point is `datasketches/tests/<sketch>_test/main.rs`,
with operation-specific modules such as `update.rs`, `union.rs`, or
`intersection.rs` alongside it.
-
-Because these entry points are nested below `tests`, Cargo does not discover
them automatically. Each new sketch target must also be registered in
`datasketches/Cargo.toml` with its required feature:
-
-```toml
-[[test]]
-name = "tuple_test"
-path = "tests/tuple_test/main.rs"
-required-features = ["tuple"]
-```
+Non-serialization tests are grouped into one integration-test target per
sketch. The target entry point is
`tests-integration/tests/<sketch>_test/main.rs`, with operation-specific
modules such as `update.rs`, `union.rs`, or `intersection.rs` alongside it.
Cargo discovers these directory-style targets automatically.
-When adding a case to an existing sketch target, add it to the appropriate
module and declare any new module from that target's `main.rs`; no Cargo
manifest change is needed. Add another `[[test]]` entry only when introducing a
new sketch target.
+When adding a case to an existing sketch target, add it to the appropriate
module and declare any new module from that target's `main.rs`. To add a sketch
target, create its directory and `main.rs`; no Cargo manifest entry or feature
gate is needed because `tests-integration` enables every sketch feature.
### Serialization compatibility tests
-Cargo automatically discovers `datasketches/tests/serde_tests.rs`, which
aggregates the sketch-specific modules under `datasketches/tests/serde_tests`.
Each module is gated by its corresponding sketch feature in `serde_tests.rs`.
+Cargo automatically discovers `tests-integration/tests/serde_tests.rs`, which
aggregates the sketch-specific modules under
`tests-integration/tests/serde_tests`.
-To add serialization tests for another sketch, add `serde_tests/<sketch>.rs`
and a feature-gated module declaration in `serde_tests.rs`. Do not add a
separate `[[test]]` entry. Shared path handling belongs in `serde_tests.rs`,
and serialization fixtures belong in the appropriate subdirectory under
`serde_tests`.
+To add serialization tests for another sketch, add `serde_tests/<sketch>.rs`
and its module declaration in `serde_tests.rs`. Shared path handling belongs in
`serde_tests.rs`, and serialization fixtures belong in the appropriate
subdirectory under `serde_tests`.
## Manual workflow (without xtask)
@@ -138,9 +137,9 @@ Serialization compatibility tests use snapshots from a
pinned revision of [`apac
The `cargo x prepare-testdata` command downloads the TCK archive and
synchronizes its snapshots into:
-- `datasketches/tests/serde_tests/cpp_generated_files`
-- `datasketches/tests/serde_tests/go_generated_files`
-- `datasketches/tests/serde_tests/java_generated_files`
+- `tests-integration/tests/serde_tests/cpp_generated_files`
+- `tests-integration/tests/serde_tests/go_generated_files`
+- `tests-integration/tests/serde_tests/java_generated_files`
You can synchronize them separately:
diff --git a/Cargo.lock b/Cargo.lock
index fe379e8..cb1b0c9 100644
--- a/Cargo.lock
+++ b/Cargo.lock
@@ -79,6 +79,14 @@ version = "0.22.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "72b3254f16251a8381aa12e40e3c4d2f0199f8c6508fbecb9d91f575e0fbb8c6"
+[[package]]
+name = "benchmarks"
+version = "0.0.0"
+dependencies = [
+ "datasketches",
+ "divan",
+]
+
[[package]]
name = "bitflags"
version = "2.13.1"
@@ -233,10 +241,8 @@ dependencies = [
name = "datasketches"
version = "0.4.0"
dependencies = [
- "divan",
"googletest",
"insta",
- "quickcheck",
"rand",
]
@@ -768,6 +774,16 @@ dependencies = [
"windows-sys 0.61.2",
]
+[[package]]
+name = "tests-integration"
+version = "0.0.0"
+dependencies = [
+ "datasketches",
+ "googletest",
+ "insta",
+ "quickcheck",
+]
+
[[package]]
name = "thiserror"
version = "2.0.19"
diff --git a/Cargo.toml b/Cargo.toml
index 46616b6..92c4775 100644
--- a/Cargo.toml
+++ b/Cargo.toml
@@ -16,7 +16,7 @@
# under the License.
[workspace]
-members = ["datasketches", "xtask"]
+members = ["benchmarks", "datasketches", "tests-integration", "xtask"]
resolver = "3"
[workspace.package]
diff --git a/benchmarks/Cargo.toml b/benchmarks/Cargo.toml
new file mode 100644
index 0000000..254de06
--- /dev/null
+++ b/benchmarks/Cargo.toml
@@ -0,0 +1,35 @@
+# Licensed to the Apache Software Foundation (ASF) under one
+# or more contributor license agreements. See the NOTICE file
+# distributed with this work for additional information
+# regarding copyright ownership. The ASF licenses this file
+# to you under the Apache License, Version 2.0 (the
+# "License"); you may not use this file except in compliance
+# with the License. You may obtain a copy of the License at
+#
+# http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing,
+# software distributed under the License is distributed on an
+# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+# KIND, either express or implied. See the License for the
+# specific language governing permissions and limitations
+# under the License.
+
+[package]
+name = "benchmarks"
+publish = false
+
+edition.workspace = true
+rust-version.workspace = true
+
+[dev-dependencies]
+datasketches = { workspace = true, features = ["cpc", "tdigest"] }
+divan = { workspace = true }
+
+[[bench]]
+harness = false
+name = "benchmarks"
+path = "main.rs"
+
+[lints]
+workspace = true
diff --git a/datasketches/tests/bloom_test/main.rs b/benchmarks/cpc/mod.rs
similarity index 98%
copy from datasketches/tests/bloom_test/main.rs
copy to benchmarks/cpc/mod.rs
index 825a628..431ad1f 100644
--- a/datasketches/tests/bloom_test/main.rs
+++ b/benchmarks/cpc/mod.rs
@@ -15,4 +15,4 @@
// specific language governing permissions and limitations
// under the License.
-mod sketch;
+mod serde;
diff --git a/benchmarks/cpc/serde.rs b/benchmarks/cpc/serde.rs
new file mode 100644
index 0000000..48b4aec
--- /dev/null
+++ b/benchmarks/cpc/serde.rs
@@ -0,0 +1,56 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use datasketches::cpc::CpcSketch;
+use divan::Bencher;
+use divan::black_box;
+use divan::black_box_drop;
+use divan::counter::BytesCount;
+use divan::counter::ItemsCount;
+
+const LG_K: u8 = 10;
+
+#[divan::bench(args = [200_u64, 8_000_u64])]
+fn serialize(bencher: Bencher, items: u64) {
+ let sketch = sketch(items);
+ let serialized_bytes = sketch.serialize().len();
+
+ bencher
+ .counter(ItemsCount::new(items))
+ .counter(BytesCount::new(serialized_bytes))
+ .bench_local(|| black_box_drop(black_box(&sketch).serialize()));
+}
+
+#[divan::bench(args = [200_u64, 8_000_u64])]
+fn deserialize(bencher: Bencher, items: u64) {
+ let bytes = sketch(items).serialize();
+
+ bencher
+ .counter(ItemsCount::new(items))
+ .counter(BytesCount::new(bytes.len()))
+ .bench_local(|| {
+ black_box_drop(CpcSketch::deserialize(black_box(&bytes)).unwrap());
+ });
+}
+
+fn sketch(items: u64) -> CpcSketch {
+ let mut sketch = CpcSketch::new(LG_K);
+ for value in 0..items {
+ sketch.update(value);
+ }
+ sketch
+}
diff --git a/datasketches/tests/bloom_test/main.rs b/benchmarks/main.rs
similarity index 83%
copy from datasketches/tests/bloom_test/main.rs
copy to benchmarks/main.rs
index 825a628..b89eebc 100644
--- a/datasketches/tests/bloom_test/main.rs
+++ b/benchmarks/main.rs
@@ -15,4 +15,14 @@
// specific language governing permissions and limitations
// under the License.
-mod sketch;
+use divan::AllocProfiler;
+
+#[global_allocator]
+static ALLOC: AllocProfiler = AllocProfiler::system();
+
+mod cpc;
+mod tdigest;
+
+fn main() {
+ divan::main();
+}
diff --git a/benchmarks/tdigest/compress.rs b/benchmarks/tdigest/compress.rs
new file mode 100644
index 0000000..3fd4cb6
--- /dev/null
+++ b/benchmarks/tdigest/compress.rs
@@ -0,0 +1,61 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use divan::Bencher;
+use divan::black_box;
+use divan::counter::ItemsCount;
+
+use super::support::build_mut_digest;
+use super::support::values;
+
+#[divan::bench]
+fn initial_buffer(bencher: Bencher) {
+ // The default k=200 digest buffers 1,640 values before automatic
compression.
+ let values = values(1_640);
+ let digest = build_mut_digest(&values);
+
+ bencher
+ .counter(ItemsCount::new(values.len()))
+ .with_inputs(|| digest.clone())
+ .bench_local_values(|mut digest| black_box(digest.rank(0.5)));
+}
+
+#[divan::bench]
+fn unmerged_tail(bencher: Bencher) {
+ let values = values(3_280);
+ let mut digest = build_mut_digest(&values[..1_640]);
+ black_box(digest.rank(0.5));
+ for &value in &values[1_640..] {
+ digest.update(value);
+ }
+
+ bencher
+ .counter(ItemsCount::new(1_640_usize))
+ .with_inputs(|| digest.clone())
+ .bench_local_values(|mut digest| black_box(digest.rank(0.5)));
+}
+
+#[divan::bench]
+fn freeze_initial_buffer(bencher: Bencher) {
+ let values = values(1_640);
+ let digest = build_mut_digest(&values);
+
+ bencher
+ .counter(ItemsCount::new(values.len()))
+ .with_inputs(|| digest.clone())
+ .bench_local_values(|digest| black_box(digest.freeze()));
+}
diff --git a/benchmarks/tdigest/merge.rs b/benchmarks/tdigest/merge.rs
new file mode 100644
index 0000000..b66b5cc
--- /dev/null
+++ b/benchmarks/tdigest/merge.rs
@@ -0,0 +1,102 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use datasketches::tdigest::TDigestMut;
+use divan::Bencher;
+use divan::black_box;
+use divan::counter::ItemsCount;
+
+use super::support::DEFAULT_DIGEST_K;
+use super::support::ROWS_PER_PARTIAL;
+use super::support::SMALL_ROWS_PER_PARTIAL;
+use super::support::build_mut_digest;
+use super::support::partial_digests;
+use super::support::partial_digests_with;
+use super::support::values;
+
+#[divan::bench]
+fn merge(bencher: Bencher) {
+ let values = values(200_000);
+ let mut left = build_mut_digest(&values[..100_000]);
+ let mut right = build_mut_digest(&values[100_000..]);
+ black_box(left.rank(0.5));
+ black_box(right.rank(0.5));
+
+ bencher
+ .counter(ItemsCount::new(values.len()))
+ .with_inputs(|| left.clone())
+ .bench_local_values(|mut left| {
+ left.merge(black_box(&right));
+ black_box(left)
+ });
+}
+
+#[divan::bench(args = [8, 1_640])]
+fn unmerged(bencher: Bencher, left_rows: usize) {
+ let left = build_mut_digest(&values(left_rows));
+ let right = build_mut_digest(&values(1_640));
+
+ bencher
+ .counter(ItemsCount::new(left_rows + 1_640))
+ .with_inputs(|| left.clone())
+ .bench_local_values(|mut left| {
+ left.merge(black_box(&right));
+ black_box(left)
+ });
+}
+
+#[divan::bench]
+fn small_partials(bencher: Bencher) {
+ let partials = partial_digests(64, SMALL_ROWS_PER_PARTIAL)
+ .into_iter()
+ .map(|mut digest| {
+ black_box(digest.rank(0.0));
+ digest
+ })
+ .collect::<Vec<_>>();
+
+ bencher
+ .counter(ItemsCount::new(64 * SMALL_ROWS_PER_PARTIAL))
+ .bench_local(|| {
+ let mut merged = TDigestMut::default();
+ for partial in &partials {
+ merged.merge(black_box(partial));
+ }
+ black_box(merged)
+ });
+}
+
+#[divan::bench]
+fn partials(bencher: Bencher) {
+ let partials = partial_digests_with(DEFAULT_DIGEST_K, 64, ROWS_PER_PARTIAL)
+ .into_iter()
+ .map(|mut digest| {
+ black_box(digest.rank(0.0));
+ digest
+ })
+ .collect::<Vec<_>>();
+
+ bencher
+ .counter(ItemsCount::new(64 * ROWS_PER_PARTIAL))
+ .bench_local(|| {
+ let mut merged = TDigestMut::new(DEFAULT_DIGEST_K);
+ for partial in &partials {
+ merged.merge(black_box(partial));
+ }
+ black_box(merged)
+ });
+}
diff --git a/datasketches/tests/hll_test/main.rs b/benchmarks/tdigest/mod.rs
similarity index 93%
copy from datasketches/tests/hll_test/main.rs
copy to benchmarks/tdigest/mod.rs
index e56cecc..adf1883 100644
--- a/datasketches/tests/hll_test/main.rs
+++ b/benchmarks/tdigest/mod.rs
@@ -15,5 +15,9 @@
// specific language governing permissions and limitations
// under the License.
-mod union;
+mod compress;
+mod merge;
+mod query;
+mod serde;
+mod support;
mod update;
diff --git a/benchmarks/tdigest/query.rs b/benchmarks/tdigest/query.rs
new file mode 100644
index 0000000..0981657
--- /dev/null
+++ b/benchmarks/tdigest/query.rs
@@ -0,0 +1,58 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use divan::Bencher;
+use divan::black_box;
+use divan::counter::ItemsCount;
+
+use super::support::prepared_digest;
+
+#[divan::bench]
+fn rank(bencher: Bencher) {
+ let digest = prepared_digest();
+
+ bencher.bench_local(|| black_box(&digest).rank(black_box(0.531_25)));
+}
+
+#[divan::bench]
+fn quantile(bencher: Bencher) {
+ let digest = prepared_digest();
+
+ bencher.bench_local(|| black_box(&digest).quantile(black_box(0.531_25)));
+}
+
+#[divan::bench]
+fn cdf_100(bencher: Bencher) {
+ let digest = prepared_digest();
+ let split_points = (1..=100).map(|i| i as f64 / 101.0).collect::<Vec<_>>();
+
+ bencher
+ .counter(ItemsCount::new(split_points.len()))
+ .bench_local(|| black_box(&digest).cdf(black_box(&split_points)));
+}
+
+#[divan::bench]
+fn quantiles_2_sequential(bencher: Bencher) {
+ let digest = prepared_digest();
+
+ bencher.bench_local(|| {
+ [
+ black_box(&digest).quantile(black_box(0.5)),
+ black_box(&digest).quantile(black_box(0.95)),
+ ]
+ });
+}
diff --git a/benchmarks/tdigest/serde.rs b/benchmarks/tdigest/serde.rs
new file mode 100644
index 0000000..6f99d04
--- /dev/null
+++ b/benchmarks/tdigest/serde.rs
@@ -0,0 +1,138 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use datasketches::tdigest::TDigestMut;
+use divan::Bencher;
+use divan::black_box;
+use divan::black_box_drop;
+use divan::counter::BytesCount;
+use divan::counter::ItemsCount;
+
+use super::support::DEFAULT_DIGEST_K;
+use super::support::PARTIAL_GROUPS;
+use super::support::ROWS_PER_PARTIAL;
+use super::support::SMALL_ROWS_PER_PARTIAL;
+use super::support::partial_digests;
+use super::support::partial_digests_with;
+use super::support::serialized_partial_digests;
+use super::support::serialized_partial_digests_with;
+use super::support::serialized_state_shape;
+use super::support::values;
+
+#[divan::bench(args = [1, 2, 8, 32, 64, 128])]
+fn small_lifecycle(bencher: Bencher, rows: usize) {
+ let values = values(rows);
+ let (serialized_bytes, _) = serialized_state_shape(DEFAULT_DIGEST_K,
&values);
+
+ bencher
+ .counter(ItemsCount::new(rows))
+ .counter(BytesCount::new(serialized_bytes))
+ .bench_local(|| {
+ let mut digest = TDigestMut::default();
+ for &value in black_box(&values) {
+ digest.update(value);
+ }
+ let bytes = digest.serialize();
+ black_box_drop(bytes);
+ black_box_drop(digest);
+ });
+}
+
+#[divan::bench(args = [10_u16, 200_u16])]
+fn partial_lifecycle_by_k(bencher: Bencher, k: u16) {
+ let values = values(ROWS_PER_PARTIAL);
+ let (serialized_bytes, centroids) = serialized_state_shape(k, &values);
+ assert!(matches!(
+ (k, serialized_bytes, centroids),
+ (10, 224, 12) | (200, 1_056, 64)
+ ));
+
+ bencher
+ .counter(ItemsCount::new(ROWS_PER_PARTIAL))
+ .counter(BytesCount::new(serialized_bytes))
+ .bench_local(|| {
+ let mut digest = TDigestMut::new(black_box(k));
+ for &value in black_box(&values) {
+ digest.update(value);
+ }
+ let bytes = digest.serialize();
+ black_box_drop(bytes);
+ black_box_drop(digest);
+ });
+}
+
+#[divan::bench]
+fn serialize_small_partial_groups(bencher: Bencher) {
+ let digests = partial_digests(PARTIAL_GROUPS, SMALL_ROWS_PER_PARTIAL);
+
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS))
+ .with_inputs(|| digests.clone())
+ .bench_local_values(|mut digests| {
+ let bytes = digests
+ .iter_mut()
+ .map(TDigestMut::serialize)
+ .collect::<Vec<_>>();
+ black_box(bytes)
+ });
+}
+
+#[divan::bench]
+fn deserialize_small_partial_groups(bencher: Bencher) {
+ let bytes = serialized_partial_digests(PARTIAL_GROUPS,
SMALL_ROWS_PER_PARTIAL);
+
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS))
+ .bench_local(|| {
+ let digests = bytes
+ .iter()
+ .map(|bytes| TDigestMut::deserialize(bytes, false).unwrap())
+ .collect::<Vec<_>>();
+ black_box(digests)
+ });
+}
+
+#[divan::bench]
+fn serialize_partial_groups(bencher: Bencher) {
+ let digests = partial_digests_with(DEFAULT_DIGEST_K, PARTIAL_GROUPS,
ROWS_PER_PARTIAL);
+
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS))
+ .with_inputs(|| digests.clone())
+ .bench_local_values(|mut digests| {
+ let bytes = digests
+ .iter_mut()
+ .map(TDigestMut::serialize)
+ .collect::<Vec<_>>();
+ black_box(bytes)
+ });
+}
+
+#[divan::bench]
+fn deserialize_partial_groups(bencher: Bencher) {
+ let bytes = serialized_partial_digests_with(DEFAULT_DIGEST_K,
PARTIAL_GROUPS, ROWS_PER_PARTIAL);
+
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS))
+ .bench_local(|| {
+ let digests = bytes
+ .iter()
+ .map(|bytes| TDigestMut::deserialize(bytes, false).unwrap())
+ .collect::<Vec<_>>();
+ black_box(digests)
+ });
+}
diff --git a/benchmarks/tdigest/support.rs b/benchmarks/tdigest/support.rs
new file mode 100644
index 0000000..64cea5a
--- /dev/null
+++ b/benchmarks/tdigest/support.rs
@@ -0,0 +1,114 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use datasketches::tdigest::TDigest;
+use datasketches::tdigest::TDigestMut;
+
+pub(super) const DEFAULT_DIGEST_K: u16 = 200;
+pub(super) const PARTIAL_GROUPS: usize = 512;
+pub(super) const ROWS_PER_PARTIAL: usize = 64;
+pub(super) const SMALL_ROWS_PER_PARTIAL: usize = 8;
+
+pub(super) fn prepared_digest() -> TDigest {
+ build_digest(&values(100_000))
+}
+
+pub(super) fn build_digest(values: &[f64]) -> TDigest {
+ build_mut_digest(values).freeze()
+}
+
+pub(super) fn build_mut_digest(values: &[f64]) -> TDigestMut {
+ let mut digest = TDigestMut::default();
+ for &value in values {
+ digest.update(value);
+ }
+ digest
+}
+
+pub(super) fn partial_digests(groups: usize, rows_per_group: usize) ->
Vec<TDigestMut> {
+ partial_digests_with(DEFAULT_DIGEST_K, groups, rows_per_group)
+}
+
+pub(super) fn partial_digests_with(
+ k: u16,
+ groups: usize,
+ rows_per_group: usize,
+) -> Vec<TDigestMut> {
+ (0..groups)
+ .map(|group| {
+ let mut digest = TDigestMut::new(k);
+ for row in 0..rows_per_group {
+ digest.update(partial_value(group, row, rows_per_group));
+ }
+ digest
+ })
+ .collect()
+}
+
+pub(super) fn serialized_partial_digests(groups: usize, rows_per_group: usize)
-> Vec<Vec<u8>> {
+ let bytes = serialized_partial_digests_with(DEFAULT_DIGEST_K, groups,
rows_per_group);
+ assert!(
+ bytes
+ .iter()
+ .all(|bytes| bytes.len() == 32 + rows_per_group * 16)
+ );
+ assert!(bytes.iter().all(|bytes| {
+ u32::from_le_bytes(bytes[8..12].try_into().unwrap()) == rows_per_group
as u32
+ }));
+ bytes
+}
+
+pub(super) fn serialized_partial_digests_with(
+ k: u16,
+ groups: usize,
+ rows_per_group: usize,
+) -> Vec<Vec<u8>> {
+ partial_digests_with(k, groups, rows_per_group)
+ .into_iter()
+ .map(|mut digest| digest.serialize())
+ .collect()
+}
+
+pub(super) fn serialized_state_shape(k: u16, values: &[f64]) -> (usize, u32) {
+ let mut digest = TDigestMut::new(k);
+ for &value in values {
+ digest.update(value);
+ }
+ let bytes = digest.serialize();
+ let centroids = match values.len() {
+ 0 => 0,
+ 1 => 1,
+ _ => u32::from_le_bytes(bytes[8..12].try_into().unwrap()),
+ };
+ (bytes.len(), centroids)
+}
+
+pub(super) fn partial_value(group: usize, row: usize, rows_per_group: usize)
-> f64 {
+ (group * rows_per_group + row) as f64
+}
+
+pub(super) fn values(len: usize) -> Vec<f64> {
+ let mut state = 0x9e37_79b9_7f4a_7c15_u64;
+ (0..len)
+ .map(|_| {
+ state ^= state << 13;
+ state ^= state >> 7;
+ state ^= state << 17;
+ (state >> 11) as f64 * (1.0 / ((1_u64 << 53) as f64))
+ })
+ .collect()
+}
diff --git a/benchmarks/tdigest/update.rs b/benchmarks/tdigest/update.rs
new file mode 100644
index 0000000..14ae2d3
--- /dev/null
+++ b/benchmarks/tdigest/update.rs
@@ -0,0 +1,115 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use datasketches::tdigest::TDigestMut;
+use divan::Bencher;
+use divan::black_box;
+use divan::counter::ItemsCount;
+
+use super::support::DEFAULT_DIGEST_K;
+use super::support::PARTIAL_GROUPS;
+use super::support::ROWS_PER_PARTIAL;
+use super::support::SMALL_ROWS_PER_PARTIAL;
+use super::support::build_digest;
+use super::support::partial_value;
+use super::support::values;
+
+#[divan::bench(args = [1_000, 100_000])]
+fn update(bencher: Bencher, len: usize) {
+ let values = values(len);
+
+ bencher
+ .counter(ItemsCount::new(len))
+ .bench_local(|| build_digest(black_box(&values)));
+}
+
+#[divan::bench]
+fn small_partial_groups_update(bencher: Bencher) {
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS * SMALL_ROWS_PER_PARTIAL))
+ .bench_local(|| {
+ let mut digests = (0..PARTIAL_GROUPS)
+ .map(|_| TDigestMut::default())
+ .collect::<Vec<_>>();
+ for (group, digest) in digests.iter_mut().enumerate() {
+ for row in 0..SMALL_ROWS_PER_PARTIAL {
+ digest.update(partial_value(group, row,
SMALL_ROWS_PER_PARTIAL));
+ }
+ }
+ black_box(digests)
+ });
+}
+
+#[divan::bench]
+fn small_partial_groups_update_two_states(bencher: Bencher) {
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS * SMALL_ROWS_PER_PARTIAL))
+ .bench_local(|| {
+ let mut digests = (0..PARTIAL_GROUPS)
+ .map(|_| (TDigestMut::default(), TDigestMut::default()))
+ .collect::<Vec<_>>();
+ for (group, (first, second)) in digests.iter_mut().enumerate() {
+ for row in 0..SMALL_ROWS_PER_PARTIAL {
+ let value = partial_value(group, row,
SMALL_ROWS_PER_PARTIAL);
+ first.update(value);
+ second.update(value);
+ }
+ }
+ black_box(digests)
+ });
+}
+
+#[divan::bench]
+fn partial_groups_update(bencher: Bencher) {
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS * ROWS_PER_PARTIAL))
+ .bench_local(|| {
+ let mut digests = (0..PARTIAL_GROUPS)
+ .map(|_| TDigestMut::new(DEFAULT_DIGEST_K))
+ .collect::<Vec<_>>();
+ for (group, digest) in digests.iter_mut().enumerate() {
+ for row in 0..ROWS_PER_PARTIAL {
+ digest.update(partial_value(group, row, ROWS_PER_PARTIAL));
+ }
+ }
+ black_box(digests)
+ });
+}
+
+#[divan::bench]
+fn partial_groups_update_two_states(bencher: Bencher) {
+ bencher
+ .counter(ItemsCount::new(PARTIAL_GROUPS * ROWS_PER_PARTIAL))
+ .bench_local(|| {
+ let mut digests = (0..PARTIAL_GROUPS)
+ .map(|_| {
+ (
+ TDigestMut::new(DEFAULT_DIGEST_K),
+ TDigestMut::new(DEFAULT_DIGEST_K),
+ )
+ })
+ .collect::<Vec<_>>();
+ for (group, (first, second)) in digests.iter_mut().enumerate() {
+ for row in 0..ROWS_PER_PARTIAL {
+ let value = partial_value(group, row, ROWS_PER_PARTIAL);
+ first.update(value);
+ second.update(value);
+ }
+ }
+ black_box(digests)
+ });
+}
diff --git a/datasketches/Cargo.toml b/datasketches/Cargo.toml
index b34d8e8..77cf206 100644
--- a/datasketches/Cargo.toml
+++ b/datasketches/Cargo.toml
@@ -48,65 +48,12 @@ tdigest = []
theta = []
tuple = []
-[[test]]
-name = "bloom_test"
-path = "tests/bloom_test/main.rs"
-required-features = ["bloom"]
-
-[[test]]
-name = "countmin_test"
-path = "tests/countmin_test/main.rs"
-required-features = ["countmin"]
-
-[[test]]
-name = "cpc_test"
-path = "tests/cpc_test/main.rs"
-required-features = ["cpc"]
-
-[[test]]
-name = "frequencies_test"
-path = "tests/frequencies_test/main.rs"
-required-features = ["frequencies"]
-
-[[test]]
-name = "hll_test"
-path = "tests/hll_test/main.rs"
-required-features = ["hll"]
-
-[[test]]
-name = "req_test"
-path = "tests/req_test/main.rs"
-required-features = ["req"]
-
-[[test]]
-name = "tdigest_test"
-path = "tests/tdigest_test/main.rs"
-required-features = ["tdigest"]
-
-[[test]]
-name = "theta_test"
-path = "tests/theta_test/main.rs"
-required-features = ["theta"]
-
-[[test]]
-name = "tuple_test"
-path = "tests/tuple_test/main.rs"
-required-features = ["tuple"]
-
-[[bench]]
-name = "tdigest"
-path = "benches/tdigest.rs"
-required-features = ["tdigest"]
-harness = false
-
[dependencies]
rand = { workspace = true, optional = true }
[dev-dependencies]
-divan = { workspace = true }
googletest = { workspace = true }
insta = { workspace = true }
-quickcheck = { workspace = true }
[lints]
workspace = true
diff --git a/datasketches/benches/tdigest.rs b/datasketches/benches/tdigest.rs
deleted file mode 100644
index b291770..0000000
--- a/datasketches/benches/tdigest.rs
+++ /dev/null
@@ -1,456 +0,0 @@
-// Licensed to the Apache Software Foundation (ASF) under one
-// or more contributor license agreements. See the NOTICE file
-// distributed with this work for additional information
-// regarding copyright ownership. The ASF licenses this file
-// to you under the Apache License, Version 2.0 (the
-// "License"); you may not use this file except in compliance
-// with the License. You may obtain a copy of the License at
-//
-// http://www.apache.org/licenses/LICENSE-2.0
-//
-// Unless required by applicable law or agreed to in writing,
-// software distributed under the License is distributed on an
-// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-// KIND, either express or implied. See the License for the
-// specific language governing permissions and limitations
-// under the License.
-
-use datasketches::tdigest::TDigest;
-use datasketches::tdigest::TDigestMut;
-use divan::AllocProfiler;
-use divan::Bencher;
-use divan::black_box;
-use divan::black_box_drop;
-use divan::counter::BytesCount;
-use divan::counter::ItemsCount;
-
-#[global_allocator]
-static ALLOC: AllocProfiler = AllocProfiler::system();
-
-const PARTIAL_GROUPS: usize = 512;
-const SMALL_ROWS_PER_PARTIAL: usize = 8;
-const DEFAULT_DIGEST_K: u16 = 200;
-const ROWS_PER_PARTIAL: usize = 64;
-
-fn main() {
- divan::main();
-}
-
-#[divan::bench(args = [1_000, 100_000])]
-fn update(bencher: Bencher, len: usize) {
- let values = values(len);
-
- bencher
- .counter(ItemsCount::new(len))
- .bench_local(|| build_digest(black_box(&values)));
-}
-
-#[divan::bench(args = [1, 2, 8, 32, 64, 128])]
-fn small_digest_lifecycle(bencher: Bencher, rows: usize) {
- let values = values(rows);
- let (serialized_bytes, _) = serialized_state_shape(200, &values);
-
- bencher
- .counter(ItemsCount::new(rows))
- .counter(BytesCount::new(serialized_bytes))
- .bench_local(|| {
- let mut digest = TDigestMut::default();
- for &value in black_box(&values) {
- digest.update(value);
- }
- let bytes = digest.serialize();
- black_box_drop(bytes);
- black_box_drop(digest);
- });
-}
-
-#[divan::bench(args = [10_u16, 200_u16])]
-fn partial_digest_lifecycle_by_k(bencher: Bencher, k: u16) {
- let values = values(ROWS_PER_PARTIAL);
- let (serialized_bytes, centroids) = serialized_state_shape(k, &values);
- assert!(matches!(
- (k, serialized_bytes, centroids),
- (10, 224, 12) | (200, 1_056, 64)
- ));
-
- bencher
- .counter(ItemsCount::new(ROWS_PER_PARTIAL))
- .counter(BytesCount::new(serialized_bytes))
- .bench_local(|| {
- let mut digest = TDigestMut::new(black_box(k));
- for &value in black_box(&values) {
- digest.update(value);
- }
- let bytes = digest.serialize();
- black_box_drop(bytes);
- black_box_drop(digest);
- });
-}
-
-#[divan::bench]
-fn compress_initial_buffer(bencher: Bencher) {
- // The default k=200 digest buffers 1,640 values before automatic
compression.
- let values = values(1_640);
- let digest = build_mut_digest(&values);
-
- bencher
- .counter(ItemsCount::new(values.len()))
- .with_inputs(|| digest.clone())
- .bench_local_values(|mut digest| black_box(digest.rank(0.5)));
-}
-
-#[divan::bench]
-fn compress_unmerged_tail(bencher: Bencher) {
- let values = values(3_280);
- let mut digest = build_mut_digest(&values[..1_640]);
- black_box(digest.rank(0.5));
- for &value in &values[1_640..] {
- digest.update(value);
- }
-
- bencher
- .counter(ItemsCount::new(1_640_usize))
- .with_inputs(|| digest.clone())
- .bench_local_values(|mut digest| black_box(digest.rank(0.5)));
-}
-
-#[divan::bench]
-fn freeze_initial_buffer(bencher: Bencher) {
- let values = values(1_640);
- let digest = build_mut_digest(&values);
-
- bencher
- .counter(ItemsCount::new(values.len()))
- .with_inputs(|| digest.clone())
- .bench_local_values(|digest| black_box(digest.freeze()));
-}
-
-#[divan::bench]
-fn merge(bencher: Bencher) {
- let values = values(200_000);
- let mut left = build_mut_digest(&values[..100_000]);
- let mut right = build_mut_digest(&values[100_000..]);
- black_box(left.rank(0.5));
- black_box(right.rank(0.5));
-
- bencher
- .counter(ItemsCount::new(values.len()))
- .with_inputs(|| left.clone())
- .bench_local_values(|mut left| {
- left.merge(black_box(&right));
- black_box(left)
- });
-}
-
-#[divan::bench(args = [8, 1_640])]
-fn unmerged_merge(bencher: Bencher, left_rows: usize) {
- let left = build_mut_digest(&values(left_rows));
- let right = build_mut_digest(&values(1_640));
-
- bencher
- .counter(ItemsCount::new(left_rows + 1_640))
- .with_inputs(|| left.clone())
- .bench_local_values(|mut left| {
- left.merge(black_box(&right));
- black_box(left)
- });
-}
-
-#[divan::bench]
-fn rank(bencher: Bencher) {
- let digest = prepared_digest();
-
- bencher.bench_local(|| black_box(&digest).rank(black_box(0.531_25)));
-}
-
-#[divan::bench]
-fn quantile(bencher: Bencher) {
- let digest = prepared_digest();
-
- bencher.bench_local(|| black_box(&digest).quantile(black_box(0.531_25)));
-}
-
-#[divan::bench]
-fn cdf_100(bencher: Bencher) {
- let digest = prepared_digest();
- let split_points = (1..=100).map(|i| i as f64 / 101.0).collect::<Vec<_>>();
-
- bencher
- .counter(ItemsCount::new(split_points.len()))
- .bench_local(|| black_box(&digest).cdf(black_box(&split_points)));
-}
-
-#[divan::bench]
-fn quantiles_2_sequential(bencher: Bencher) {
- let digest = prepared_digest();
-
- bencher.bench_local(|| {
- [
- black_box(&digest).quantile(black_box(0.5)),
- black_box(&digest).quantile(black_box(0.95)),
- ]
- });
-}
-
-#[divan::bench]
-fn small_partial_groups_update(bencher: Bencher) {
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS * SMALL_ROWS_PER_PARTIAL))
- .bench_local(|| {
- let mut digests = (0..PARTIAL_GROUPS)
- .map(|_| TDigestMut::default())
- .collect::<Vec<_>>();
- for (group, digest) in digests.iter_mut().enumerate() {
- for row in 0..SMALL_ROWS_PER_PARTIAL {
- digest.update(partial_value(group, row,
SMALL_ROWS_PER_PARTIAL));
- }
- }
- black_box(digests)
- });
-}
-
-#[divan::bench]
-fn small_partial_groups_update_two_states(bencher: Bencher) {
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS * SMALL_ROWS_PER_PARTIAL))
- .bench_local(|| {
- let mut digests = (0..PARTIAL_GROUPS)
- .map(|_| (TDigestMut::default(), TDigestMut::default()))
- .collect::<Vec<_>>();
- for (group, (first, second)) in digests.iter_mut().enumerate() {
- for row in 0..SMALL_ROWS_PER_PARTIAL {
- let value = partial_value(group, row,
SMALL_ROWS_PER_PARTIAL);
- first.update(value);
- second.update(value);
- }
- }
- black_box(digests)
- });
-}
-
-#[divan::bench]
-fn small_partial_groups_serialize(bencher: Bencher) {
- let digests = partial_digests(PARTIAL_GROUPS, SMALL_ROWS_PER_PARTIAL);
-
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS))
- .with_inputs(|| digests.clone())
- .bench_local_values(|mut digests| {
- let bytes = digests
- .iter_mut()
- .map(TDigestMut::serialize)
- .collect::<Vec<_>>();
- black_box(bytes)
- });
-}
-
-#[divan::bench]
-fn small_partial_groups_deserialize(bencher: Bencher) {
- let bytes = serialized_partial_digests(PARTIAL_GROUPS,
SMALL_ROWS_PER_PARTIAL);
-
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS))
- .bench_local(|| {
- let digests = bytes
- .iter()
- .map(|bytes| TDigestMut::deserialize(bytes, false).unwrap())
- .collect::<Vec<_>>();
- black_box(digests)
- });
-}
-
-#[divan::bench]
-fn small_partial_merge(bencher: Bencher) {
- let partials = partial_digests(64, SMALL_ROWS_PER_PARTIAL)
- .into_iter()
- .map(|mut digest| {
- black_box(digest.rank(0.0));
- digest
- })
- .collect::<Vec<_>>();
-
- bencher
- .counter(ItemsCount::new(64 * SMALL_ROWS_PER_PARTIAL))
- .bench_local(|| {
- let mut merged = TDigestMut::default();
- for partial in &partials {
- merged.merge(black_box(partial));
- }
- black_box(merged)
- });
-}
-
-#[divan::bench]
-fn partial_groups_update(bencher: Bencher) {
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS * ROWS_PER_PARTIAL))
- .bench_local(|| {
- let mut digests = (0..PARTIAL_GROUPS)
- .map(|_| TDigestMut::new(DEFAULT_DIGEST_K))
- .collect::<Vec<_>>();
- for (group, digest) in digests.iter_mut().enumerate() {
- for row in 0..ROWS_PER_PARTIAL {
- digest.update(partial_value(group, row, ROWS_PER_PARTIAL));
- }
- }
- black_box(digests)
- });
-}
-
-#[divan::bench]
-fn partial_groups_update_two_states(bencher: Bencher) {
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS * ROWS_PER_PARTIAL))
- .bench_local(|| {
- let mut digests = (0..PARTIAL_GROUPS)
- .map(|_| {
- (
- TDigestMut::new(DEFAULT_DIGEST_K),
- TDigestMut::new(DEFAULT_DIGEST_K),
- )
- })
- .collect::<Vec<_>>();
- for (group, (first, second)) in digests.iter_mut().enumerate() {
- for row in 0..ROWS_PER_PARTIAL {
- let value = partial_value(group, row, ROWS_PER_PARTIAL);
- first.update(value);
- second.update(value);
- }
- }
- black_box(digests)
- });
-}
-
-#[divan::bench]
-fn partial_groups_serialize(bencher: Bencher) {
- let digests = partial_digests_with(DEFAULT_DIGEST_K, PARTIAL_GROUPS,
ROWS_PER_PARTIAL);
-
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS))
- .with_inputs(|| digests.clone())
- .bench_local_values(|mut digests| {
- let bytes = digests
- .iter_mut()
- .map(TDigestMut::serialize)
- .collect::<Vec<_>>();
- black_box(bytes)
- });
-}
-
-#[divan::bench]
-fn partial_groups_deserialize(bencher: Bencher) {
- let bytes = serialized_partial_digests_with(DEFAULT_DIGEST_K,
PARTIAL_GROUPS, ROWS_PER_PARTIAL);
-
- bencher
- .counter(ItemsCount::new(PARTIAL_GROUPS))
- .bench_local(|| {
- let digests = bytes
- .iter()
- .map(|bytes| TDigestMut::deserialize(bytes, false).unwrap())
- .collect::<Vec<_>>();
- black_box(digests)
- });
-}
-
-#[divan::bench]
-fn partial_merge(bencher: Bencher) {
- let partials = partial_digests_with(DEFAULT_DIGEST_K, 64, ROWS_PER_PARTIAL)
- .into_iter()
- .map(|mut digest| {
- black_box(digest.rank(0.0));
- digest
- })
- .collect::<Vec<_>>();
-
- bencher
- .counter(ItemsCount::new(64 * ROWS_PER_PARTIAL))
- .bench_local(|| {
- let mut merged = TDigestMut::new(DEFAULT_DIGEST_K);
- for partial in &partials {
- merged.merge(black_box(partial));
- }
- black_box(merged)
- });
-}
-
-fn prepared_digest() -> TDigest {
- build_digest(&values(100_000))
-}
-
-fn build_digest(values: &[f64]) -> TDigest {
- build_mut_digest(values).freeze()
-}
-
-fn build_mut_digest(values: &[f64]) -> TDigestMut {
- let mut digest = TDigestMut::default();
- for &value in values {
- digest.update(value);
- }
- digest
-}
-
-fn partial_digests(groups: usize, rows_per_group: usize) -> Vec<TDigestMut> {
- partial_digests_with(200, groups, rows_per_group)
-}
-
-fn partial_digests_with(k: u16, groups: usize, rows_per_group: usize) ->
Vec<TDigestMut> {
- (0..groups)
- .map(|group| {
- let mut digest = TDigestMut::new(k);
- for row in 0..rows_per_group {
- digest.update(partial_value(group, row, rows_per_group));
- }
- digest
- })
- .collect()
-}
-
-fn serialized_partial_digests(groups: usize, rows_per_group: usize) ->
Vec<Vec<u8>> {
- let bytes = serialized_partial_digests_with(200, groups, rows_per_group);
- assert!(
- bytes
- .iter()
- .all(|bytes| bytes.len() == 32 + rows_per_group * 16)
- );
- assert!(bytes.iter().all(|bytes| {
- u32::from_le_bytes(bytes[8..12].try_into().unwrap()) == rows_per_group
as u32
- }));
- bytes
-}
-
-fn serialized_partial_digests_with(k: u16, groups: usize, rows_per_group:
usize) -> Vec<Vec<u8>> {
- partial_digests_with(k, groups, rows_per_group)
- .into_iter()
- .map(|mut digest| digest.serialize())
- .collect()
-}
-
-fn serialized_state_shape(k: u16, values: &[f64]) -> (usize, u32) {
- let mut digest = TDigestMut::new(k);
- for &value in values {
- digest.update(value);
- }
- let bytes = digest.serialize();
- let centroids = match values.len() {
- 0 => 0,
- 1 => 1,
- _ => u32::from_le_bytes(bytes[8..12].try_into().unwrap()),
- };
- (bytes.len(), centroids)
-}
-
-fn partial_value(group: usize, row: usize, rows_per_group: usize) -> f64 {
- (group * rows_per_group + row) as f64
-}
-
-fn values(len: usize) -> Vec<f64> {
- let mut state = 0x9e37_79b9_7f4a_7c15_u64;
- (0..len)
- .map(|_| {
- state ^= state << 13;
- state ^= state >> 7;
- state ^= state << 17;
- (state >> 11) as f64 * (1.0 / ((1_u64 << 53) as f64))
- })
- .collect()
-}
diff --git a/tests-integration/Cargo.toml b/tests-integration/Cargo.toml
new file mode 100644
index 0000000..0e1f808
--- /dev/null
+++ b/tests-integration/Cargo.toml
@@ -0,0 +1,48 @@
+# Licensed to the Apache Software Foundation (ASF) under one
+# or more contributor license agreements. See the NOTICE file
+# distributed with this work for additional information
+# regarding copyright ownership. The ASF licenses this file
+# to you under the Apache License, Version 2.0 (the
+# "License"); you may not use this file except in compliance
+# with the License. You may obtain a copy of the License at
+#
+# http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing,
+# software distributed under the License is distributed on an
+# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+# KIND, either express or implied. See the License for the
+# specific language governing permissions and limitations
+# under the License.
+
+[package]
+name = "tests-integration"
+publish = false
+
+edition.workspace = true
+rust-version.workspace = true
+
+[dev-dependencies]
+datasketches = { workspace = true, features = [
+ "bloom",
+ "countmin",
+ "cpc",
+ "frequencies",
+ "hll",
+ "req",
+ "tdigest",
+ "theta",
+ "tuple",
+] }
+googletest = { workspace = true }
+insta = { workspace = true }
+quickcheck = { workspace = true }
+
+[lib]
+bench = false
+doc = false
+doctest = false
+test = false
+
+[lints]
+workspace = true
diff --git a/datasketches/tests/bloom_test/main.rs
b/tests-integration/src/lib.rs
similarity index 93%
copy from datasketches/tests/bloom_test/main.rs
copy to tests-integration/src/lib.rs
index 825a628..4358913 100644
--- a/datasketches/tests/bloom_test/main.rs
+++ b/tests-integration/src/lib.rs
@@ -15,4 +15,4 @@
// specific language governing permissions and limitations
// under the License.
-mod sketch;
+//! Shared support for end-to-end integration tests.
diff --git a/datasketches/tests/bloom_test/main.rs
b/tests-integration/tests/bloom_test/main.rs
similarity index 100%
rename from datasketches/tests/bloom_test/main.rs
rename to tests-integration/tests/bloom_test/main.rs
diff --git a/datasketches/tests/bloom_test/sketch.rs
b/tests-integration/tests/bloom_test/sketch.rs
similarity index 100%
rename from datasketches/tests/bloom_test/sketch.rs
rename to tests-integration/tests/bloom_test/sketch.rs
diff --git a/datasketches/tests/countmin_test/main.rs
b/tests-integration/tests/countmin_test/main.rs
similarity index 100%
rename from datasketches/tests/countmin_test/main.rs
rename to tests-integration/tests/countmin_test/main.rs
diff --git a/datasketches/tests/countmin_test/sketch.rs
b/tests-integration/tests/countmin_test/sketch.rs
similarity index 100%
rename from datasketches/tests/countmin_test/sketch.rs
rename to tests-integration/tests/countmin_test/sketch.rs
diff --git a/datasketches/tests/cpc_test/deserialize.rs
b/tests-integration/tests/cpc_test/deserialize.rs
similarity index 100%
rename from datasketches/tests/cpc_test/deserialize.rs
rename to tests-integration/tests/cpc_test/deserialize.rs
diff --git a/datasketches/tests/cpc_test/main.rs
b/tests-integration/tests/cpc_test/main.rs
similarity index 100%
rename from datasketches/tests/cpc_test/main.rs
rename to tests-integration/tests/cpc_test/main.rs
diff --git a/datasketches/tests/cpc_test/union.rs
b/tests-integration/tests/cpc_test/union.rs
similarity index 100%
rename from datasketches/tests/cpc_test/union.rs
rename to tests-integration/tests/cpc_test/union.rs
diff --git a/datasketches/tests/cpc_test/update.rs
b/tests-integration/tests/cpc_test/update.rs
similarity index 100%
rename from datasketches/tests/cpc_test/update.rs
rename to tests-integration/tests/cpc_test/update.rs
diff --git a/datasketches/tests/cpc_test/wrapper.rs
b/tests-integration/tests/cpc_test/wrapper.rs
similarity index 100%
rename from datasketches/tests/cpc_test/wrapper.rs
rename to tests-integration/tests/cpc_test/wrapper.rs
diff --git a/datasketches/tests/frequencies_test/main.rs
b/tests-integration/tests/frequencies_test/main.rs
similarity index 100%
rename from datasketches/tests/frequencies_test/main.rs
rename to tests-integration/tests/frequencies_test/main.rs
diff --git a/datasketches/tests/frequencies_test/update.rs
b/tests-integration/tests/frequencies_test/update.rs
similarity index 100%
rename from datasketches/tests/frequencies_test/update.rs
rename to tests-integration/tests/frequencies_test/update.rs
diff --git a/datasketches/tests/hll_test/main.rs
b/tests-integration/tests/hll_test/main.rs
similarity index 100%
rename from datasketches/tests/hll_test/main.rs
rename to tests-integration/tests/hll_test/main.rs
diff --git a/datasketches/tests/hll_test/union.rs
b/tests-integration/tests/hll_test/union.rs
similarity index 100%
rename from datasketches/tests/hll_test/union.rs
rename to tests-integration/tests/hll_test/union.rs
diff --git a/datasketches/tests/hll_test/update.rs
b/tests-integration/tests/hll_test/update.rs
similarity index 100%
rename from datasketches/tests/hll_test/update.rs
rename to tests-integration/tests/hll_test/update.rs
diff --git a/datasketches/tests/req_test/accuracy.rs
b/tests-integration/tests/req_test/accuracy.rs
similarity index 100%
rename from datasketches/tests/req_test/accuracy.rs
rename to tests-integration/tests/req_test/accuracy.rs
diff --git a/datasketches/tests/req_test/bounds.rs
b/tests-integration/tests/req_test/bounds.rs
similarity index 100%
rename from datasketches/tests/req_test/bounds.rs
rename to tests-integration/tests/req_test/bounds.rs
diff --git a/datasketches/tests/req_test/core.rs
b/tests-integration/tests/req_test/core.rs
similarity index 100%
rename from datasketches/tests/req_test/core.rs
rename to tests-integration/tests/req_test/core.rs
diff --git a/datasketches/tests/req_test/main.rs
b/tests-integration/tests/req_test/main.rs
similarity index 100%
rename from datasketches/tests/req_test/main.rs
rename to tests-integration/tests/req_test/main.rs
diff --git a/datasketches/tests/req_test/merge.rs
b/tests-integration/tests/req_test/merge.rs
similarity index 100%
rename from datasketches/tests/req_test/merge.rs
rename to tests-integration/tests/req_test/merge.rs
diff --git a/datasketches/tests/req_test/property.rs
b/tests-integration/tests/req_test/property.rs
similarity index 100%
rename from datasketches/tests/req_test/property.rs
rename to tests-integration/tests/req_test/property.rs
diff --git a/datasketches/tests/req_test/query.rs
b/tests-integration/tests/req_test/query.rs
similarity index 100%
rename from datasketches/tests/req_test/query.rs
rename to tests-integration/tests/req_test/query.rs
diff --git a/datasketches/tests/req_test/sorted_view_api.rs
b/tests-integration/tests/req_test/sorted_view_api.rs
similarity index 100%
rename from datasketches/tests/req_test/sorted_view_api.rs
rename to tests-integration/tests/req_test/sorted_view_api.rs
diff --git a/datasketches/tests/req_test/structure.rs
b/tests-integration/tests/req_test/structure.rs
similarity index 100%
rename from datasketches/tests/req_test/structure.rs
rename to tests-integration/tests/req_test/structure.rs
diff --git a/datasketches/tests/req_test/union.rs
b/tests-integration/tests/req_test/union.rs
similarity index 100%
rename from datasketches/tests/req_test/union.rs
rename to tests-integration/tests/req_test/union.rs
diff --git a/datasketches/tests/serde_tests.rs
b/tests-integration/tests/serde_tests.rs
similarity index 88%
rename from datasketches/tests/serde_tests.rs
rename to tests-integration/tests/serde_tests.rs
index 5de4587..7120877 100644
--- a/datasketches/tests/serde_tests.rs
+++ b/tests-integration/tests/serde_tests.rs
@@ -42,38 +42,29 @@ pub fn serialization_test_data(sub_dir: &str, name: &str)
-> PathBuf {
path
}
-#[cfg(feature = "bloom")]
#[path = "serde_tests/bloom.rs"]
mod bloom;
-#[cfg(feature = "countmin")]
#[path = "serde_tests/countmin.rs"]
mod countmin;
-#[cfg(feature = "cpc")]
#[path = "serde_tests/cpc.rs"]
mod cpc;
-#[cfg(feature = "frequencies")]
#[path = "serde_tests/frequencies.rs"]
mod frequencies;
-#[cfg(feature = "hll")]
#[path = "serde_tests/hll.rs"]
mod hll;
-#[cfg(feature = "req")]
#[path = "serde_tests/req.rs"]
mod req;
-#[cfg(feature = "tdigest")]
#[path = "serde_tests/tdigest.rs"]
mod tdigest;
-#[cfg(feature = "theta")]
#[path = "serde_tests/theta.rs"]
mod theta;
-#[cfg(feature = "tuple")]
#[path = "serde_tests/tuple.rs"]
mod tuple;
diff --git a/datasketches/tests/serde_tests/.gitignore
b/tests-integration/tests/serde_tests/.gitignore
similarity index 100%
rename from datasketches/tests/serde_tests/.gitignore
rename to tests-integration/tests/serde_tests/.gitignore
diff --git a/datasketches/tests/serde_tests/bloom.rs
b/tests-integration/tests/serde_tests/bloom.rs
similarity index 100%
rename from datasketches/tests/serde_tests/bloom.rs
rename to tests-integration/tests/serde_tests/bloom.rs
diff --git a/datasketches/tests/serde_tests/countmin.rs
b/tests-integration/tests/serde_tests/countmin.rs
similarity index 100%
rename from datasketches/tests/serde_tests/countmin.rs
rename to tests-integration/tests/serde_tests/countmin.rs
diff --git a/datasketches/tests/serde_tests/cpc.rs
b/tests-integration/tests/serde_tests/cpc.rs
similarity index 100%
rename from datasketches/tests/serde_tests/cpc.rs
rename to tests-integration/tests/serde_tests/cpc.rs
diff --git a/datasketches/tests/serde_tests/frequencies.rs
b/tests-integration/tests/serde_tests/frequencies.rs
similarity index 100%
rename from datasketches/tests/serde_tests/frequencies.rs
rename to tests-integration/tests/serde_tests/frequencies.rs
diff --git a/datasketches/tests/serde_tests/hll.rs
b/tests-integration/tests/serde_tests/hll.rs
similarity index 100%
rename from datasketches/tests/serde_tests/hll.rs
rename to tests-integration/tests/serde_tests/hll.rs
diff --git
a/datasketches/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_double.sk
b/tests-integration/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_double.sk
similarity index 100%
rename from
datasketches/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_double.sk
rename to
tests-integration/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_double.sk
diff --git
a/datasketches/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_float.sk
b/tests-integration/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_float.sk
similarity index 100%
rename from
datasketches/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_float.sk
rename to
tests-integration/tests/serde_tests/reference_files/tdigest_ref_k100_n10000_float.sk
diff --git a/datasketches/tests/serde_tests/req.rs
b/tests-integration/tests/serde_tests/req.rs
similarity index 100%
rename from datasketches/tests/serde_tests/req.rs
rename to tests-integration/tests/serde_tests/req.rs
diff --git a/datasketches/tests/serde_tests/tdigest.rs
b/tests-integration/tests/serde_tests/tdigest.rs
similarity index 100%
rename from datasketches/tests/serde_tests/tdigest.rs
rename to tests-integration/tests/serde_tests/tdigest.rs
diff --git a/datasketches/tests/serde_tests/theta.rs
b/tests-integration/tests/serde_tests/theta.rs
similarity index 100%
rename from datasketches/tests/serde_tests/theta.rs
rename to tests-integration/tests/serde_tests/theta.rs
diff --git a/datasketches/tests/serde_tests/tuple.rs
b/tests-integration/tests/serde_tests/tuple.rs
similarity index 100%
rename from datasketches/tests/serde_tests/tuple.rs
rename to tests-integration/tests/serde_tests/tuple.rs
diff --git a/datasketches/tests/tdigest_test/main.rs
b/tests-integration/tests/tdigest_test/main.rs
similarity index 100%
rename from datasketches/tests/tdigest_test/main.rs
rename to tests-integration/tests/tdigest_test/main.rs
diff --git a/datasketches/tests/tdigest_test/sketch.rs
b/tests-integration/tests/tdigest_test/sketch.rs
similarity index 100%
rename from datasketches/tests/tdigest_test/sketch.rs
rename to tests-integration/tests/tdigest_test/sketch.rs
diff --git a/datasketches/tests/theta_test/a_not_b.rs
b/tests-integration/tests/theta_test/a_not_b.rs
similarity index 100%
rename from datasketches/tests/theta_test/a_not_b.rs
rename to tests-integration/tests/theta_test/a_not_b.rs
diff --git a/datasketches/tests/theta_test/intersection.rs
b/tests-integration/tests/theta_test/intersection.rs
similarity index 100%
rename from datasketches/tests/theta_test/intersection.rs
rename to tests-integration/tests/theta_test/intersection.rs
diff --git a/datasketches/tests/theta_test/jaccard_similarity.rs
b/tests-integration/tests/theta_test/jaccard_similarity.rs
similarity index 100%
rename from datasketches/tests/theta_test/jaccard_similarity.rs
rename to tests-integration/tests/theta_test/jaccard_similarity.rs
diff --git a/datasketches/tests/theta_test/main.rs
b/tests-integration/tests/theta_test/main.rs
similarity index 100%
rename from datasketches/tests/theta_test/main.rs
rename to tests-integration/tests/theta_test/main.rs
diff --git a/datasketches/tests/theta_test/sketch.rs
b/tests-integration/tests/theta_test/sketch.rs
similarity index 100%
rename from datasketches/tests/theta_test/sketch.rs
rename to tests-integration/tests/theta_test/sketch.rs
diff --git a/datasketches/tests/theta_test/union.rs
b/tests-integration/tests/theta_test/union.rs
similarity index 100%
rename from datasketches/tests/theta_test/union.rs
rename to tests-integration/tests/theta_test/union.rs
diff --git a/datasketches/tests/tuple_test/a_not_b.rs
b/tests-integration/tests/tuple_test/a_not_b.rs
similarity index 100%
rename from datasketches/tests/tuple_test/a_not_b.rs
rename to tests-integration/tests/tuple_test/a_not_b.rs
diff --git a/datasketches/tests/tuple_test/intersection.rs
b/tests-integration/tests/tuple_test/intersection.rs
similarity index 100%
rename from datasketches/tests/tuple_test/intersection.rs
rename to tests-integration/tests/tuple_test/intersection.rs
diff --git a/datasketches/tests/tuple_test/jaccard_similarity.rs
b/tests-integration/tests/tuple_test/jaccard_similarity.rs
similarity index 100%
rename from datasketches/tests/tuple_test/jaccard_similarity.rs
rename to tests-integration/tests/tuple_test/jaccard_similarity.rs
diff --git a/datasketches/tests/tuple_test/main.rs
b/tests-integration/tests/tuple_test/main.rs
similarity index 100%
rename from datasketches/tests/tuple_test/main.rs
rename to tests-integration/tests/tuple_test/main.rs
diff --git a/datasketches/tests/tuple_test/sketch.rs
b/tests-integration/tests/tuple_test/sketch.rs
similarity index 100%
rename from datasketches/tests/tuple_test/sketch.rs
rename to tests-integration/tests/tuple_test/sketch.rs
diff --git a/datasketches/tests/tuple_test/union.rs
b/tests-integration/tests/tuple_test/union.rs
similarity index 100%
rename from datasketches/tests/tuple_test/union.rs
rename to tests-integration/tests/tuple_test/union.rs
diff --git a/xtask/src/main.rs b/xtask/src/main.rs
index bac79cd..9db910e 100644
--- a/xtask/src/main.rs
+++ b/xtask/src/main.rs
@@ -192,7 +192,7 @@ fn run_command(mut cmd: StdCommand) {
fn make_bench_cmd() -> StdCommand {
let mut cmd = find_command("cargo");
- cmd.args(["bench", "--workspace", "--all-features", "--bench", "*"]);
+ cmd.args(["bench", "--package", "benchmarks", "--bench", "benchmarks"]);
cmd
}
@@ -318,7 +318,7 @@ impl CommandPrepareTestData {
fn prepare(self) -> Result<()> {
const REVISION: &str = "d363b12d293b395d90abb42677f9ea63178dbc0d";
let serde_tests =
-
Path::new(env!("CARGO_WORKSPACE_DIR")).join("datasketches/tests/serde_tests");
+
Path::new(env!("CARGO_WORKSPACE_DIR")).join("tests-integration/tests/serde_tests");
let archive_url =
format!("https://api.github.com/repos/apache/datasketches-tck/tarball/{REVISION}");
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]