Hi Peter, Thanks for the context and the detailed explanation! It helps a lot in clarifying the motivation.
I still have some questions about the second alternative you mentioned: *bounded overlay files (mini-LSM style), with fixed ranges and one or more update files.* Specifically: - Why would the read time and resolution time still be unpredictable? For example, if we allow at most 5 region files for a given range, then the worst case is still bounded: read up to 5 files and reconcile the results. In addition, if we also store per-region-file statistics or Bloom filters in the tracking file, I would expect the worst case to happen much less frequently. - I understand that once the file-count limit is reached, the maintenance service would eventually need to compact those files back down. But that still seems useful because it gives the maintainer some buffer to absorb multiple small updates before rewriting the base region, rather than requiring every update to rewrite the affected region immediately. - For covering indexes in particular, it also seems that spending a bit more time reading a few index files could still be a good trade-off, since we can avoid going back to the base table entirely. So I'm wondering whether the concern here is mainly the additional object-store reads, the reconciliation cost, or something else that makes the latency difficult to bound even when the number of overlay files is strictly limited. Best, Shawn
