leerho commented on PR #515:
URL: https://github.com/apache/datasketches-cpp/pull/515#issuecomment-5483838182
Is predictable storage space the concern? If so, why don't you just reduce
Lg_K by one? All sketches will fit in the same space as your anticipated
Compact-Trimmed sketch. The guaranteed RSE will be increased by about 41%
($sqrt(2)-1$), but over a large # of sketches the average RSE increase will be
about 20%. For example, at a LgK=14, +/-RSE is 0.78% and at LgK=13, +/-RSE is
1.1%. And over a large number of sketches what you are likely to see on
average is about half that difference.
Or is a predictable RSE distribution more important than overall improved
accuracy? That is the one case where I could see the desire to trim to K.
However, you don't need to trim the sketch size to obtain a very predictable
RSE distribution.
You can compute this special estimate yourself:
```
public double getSpecialEstimate(ThetaSketch sk, int lgK) {
int k = 1 << lgK;
double theta = sk.getTheta();
int entries = sk.getRetainedEntries();
return (entries <= k) ? entries : k / theta;
}
```
This is super fast and no rebuilding required.
If the given sketch is a CompactThetaSketch you will need to supply the
chosen LgK of the original ThetaSketch, because CompactSketches don't need K
(or LgK).
I am struggling to justify adding such a special API to a general purpose
library.
I would feel more sympathetic if you could explain what your concern is and
what you are trying to achieve by alway degrading your accuracy to the minimum
guaranteed statistical accuracy.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]