Hi all,

I would like to start a discussion on restructuring datasketches-cpp for a
6.0.0 major release. There are three related pieces: the file structure,
the namespace hierarchy, and moving the required language level from C++11
to C++17.

The proposal is written up in two documents, posted as a draft PR so they
can be commented on line by line. The PR is not intended to merge; it is
just a place to mark up the text.

  https://github.com/apache/datasketches-cpp/pull/528

  restructure-draft-v2.md   - target file structure, namespace hierarchy,
definitions, ratified decisions, open items
  restructure-plan-detail.md - the PR sequence, per-area breakdown, and how
we keep each PR reviewable

In brief, what is proposed:

* A single include root. Headers move from <sketch>/include/ to
include/datasketches/<sketch>/, so a user writes #include
<datasketches/theta/theta_sketch.hpp> instead of #include
<theta_sketch.hpp>. This ends the collisions that flat names like serde.hpp
and version.hpp invite on a shared include path.

* One nested namespace per sketch family: datasketches::theta,
datasketches::hll, and so on, with datasketches::<area>::internal for
everything outside the supported API. This mirrors the Java package layout,
and follows the same model as Boost, Apache Arrow, Apache Thrift, POCO and
the AWS C++ SDK.

* An internal/ subdirectory per area for private classes and helpers.
Because the library is header-only these headers still ship; internal means
"not covered by compatibility guarantees", not "hidden".

* C++11 -> C++17, and then adopting C++17 constructs one feature class per
PR. Nothing C++17 removed is currently in use, and every compiler in our CI
matrix already supports it.

This is a breaking change for every user: include paths and namespaces both
change, and the installed location moves from include/DataSketches/ to
include/datasketches/. That is why it is 6.0.0. datasketches-python,
-postgresql and -bigquery will each need updating, and they can validate
against 6.0.0-RC1 during the release vote.

A few specific decisions I would especially like feedback on:

* Renaming the fi/ directory to frequencies/ to match the Java package name.
* Renaming the HLL headers to snake_case, so they match every other sketch.
* Moving bounds_binomial_proportions, bounds_on_ratios_in_sampled_sets and
bounds_on_ratios_in_theta_sketched_sets into internal/. These are public in
Java only so its tests can reach them across packages; they are probability
math, not an API users would call directly.
* Sub-namespaces for the category directories: sampling::var_opt,
sampling::ebpps, filters::bloom, tuple::array_tuple, tuple::aod, tuple::aos.

Comments inline on the PR are welcome and probably the easiest way to mark
up specific lines, but please bring anything that needs a decision back to
this thread so it is recorded here.

I would like to leave this open for comment for the next two weeks, and
then summarize what we have agreed before any code moves.

I would like to close this on October 6th, unless there is ongoing active
discussion.

I will be traveling from October 11 to 24, so I would like to have the
discussion wrapped up before then; anything still open when I leave will
wait until I am back.

Thanks,
Lee

Reply via email to