Hi Peter,

Thanks for the context and the detailed explanation! It helps a lot in
clarifying the motivation.

I still have some questions about the second alternative you
mentioned: *bounded
overlay files (mini-LSM style), with fixed ranges and one or more update
files.*

Specifically:

- Why would the read time and resolution time still be unpredictable? For
example, if we allow at most 5 region files for a given range, then the
worst case is still bounded: read up to 5 files and reconcile the results.
In addition, if we also store per-region-file statistics or Bloom filters
in the tracking file, I would expect the worst case to happen much less
frequently.

- I understand that once the file-count limit is reached, the maintenance
service would eventually need to compact those files back down. But that
still seems useful because it gives the maintainer some buffer to absorb
multiple small updates before rewriting the base region, rather than
requiring every update to rewrite the affected region immediately.

- For covering indexes in particular, it also seems that spending a bit
more time reading a few index files could still be a good trade-off, since
we can avoid going back to the base table entirely.

So I'm wondering whether the concern here is mainly the additional
object-store reads, the reconciliation cost, or something else that makes
the latency difficult to bound even when the number of overlay files is
strictly limited.


Best,

Shawn

Reply via email to