Hi all, I would like to start a discussion on restructuring datasketches-cpp for a 6.0.0 major release. There are three related pieces: the file structure, the namespace hierarchy, and moving the required language level from C++11 to C++17.
The proposal is written up in two documents, posted as a draft PR so they can be commented on line by line. The PR is not intended to merge; it is just a place to mark up the text. https://github.com/apache/datasketches-cpp/pull/528 restructure-draft-v2.md - target file structure, namespace hierarchy, definitions, ratified decisions, open items restructure-plan-detail.md - the PR sequence, per-area breakdown, and how we keep each PR reviewable In brief, what is proposed: * A single include root. Headers move from <sketch>/include/ to include/datasketches/<sketch>/, so a user writes #include <datasketches/theta/theta_sketch.hpp> instead of #include <theta_sketch.hpp>. This ends the collisions that flat names like serde.hpp and version.hpp invite on a shared include path. * One nested namespace per sketch family: datasketches::theta, datasketches::hll, and so on, with datasketches::<area>::internal for everything outside the supported API. This mirrors the Java package layout, and follows the same model as Boost, Apache Arrow, Apache Thrift, POCO and the AWS C++ SDK. * An internal/ subdirectory per area for private classes and helpers. Because the library is header-only these headers still ship; internal means "not covered by compatibility guarantees", not "hidden". * C++11 -> C++17, and then adopting C++17 constructs one feature class per PR. Nothing C++17 removed is currently in use, and every compiler in our CI matrix already supports it. This is a breaking change for every user: include paths and namespaces both change, and the installed location moves from include/DataSketches/ to include/datasketches/. That is why it is 6.0.0. datasketches-python, -postgresql and -bigquery will each need updating, and they can validate against 6.0.0-RC1 during the release vote. A few specific decisions I would especially like feedback on: * Renaming the fi/ directory to frequencies/ to match the Java package name. * Renaming the HLL headers to snake_case, so they match every other sketch. * Moving bounds_binomial_proportions, bounds_on_ratios_in_sampled_sets and bounds_on_ratios_in_theta_sketched_sets into internal/. These are public in Java only so its tests can reach them across packages; they are probability math, not an API users would call directly. * Sub-namespaces for the category directories: sampling::var_opt, sampling::ebpps, filters::bloom, tuple::array_tuple, tuple::aod, tuple::aos. Comments inline on the PR are welcome and probably the easiest way to mark up specific lines, but please bring anything that needs a decision back to this thread so it is recorded here. I would like to leave this open for comment for the next two weeks, and then summarize what we have agreed before any code moves. I would like to close this on October 6th, unless there is ongoing active discussion. I will be traveling from October 11 to 24, so I would like to have the discussion wrapped up before then; anything still open when I leave will wait until I am back. Thanks, Lee
