Hi, Andrey
Thanks for updating the patches. On Fri, 07 Aug 2026 at 18:15, Andrey Borodin <[email protected]> wrote: >> On 26 Jul 2026, at 21:09, Andrey Borodin <[email protected]> wrote: >> >> This is v7 > > Hi hackers! > > This is v8. Still three patches. 0003 now answers the question I left > open in v7 - a reader can start at an arbitrary LSN. 0001 grew to > cover decompression as well, which turned out to be the bigger win, and > running the test suites with the feature actually turned on found a > number of bugs. > > Numbers are from two machines, 4-core for 0002 and the compression side > of 0001, 16-core for 0003 and the redo figures, both built with -O2 and > without assertions, fsync off, shared_buffers 8GB. Each was repeated, > where a number moved between runs I say so. > > > ---- 0001: reuse zstd contexts, both directions ---- > > Still the piece that is useful on its own, and it grew since v7: it now > keeps the decompression context too, which turns out to matter more. > > XLogCompressBackupBlock() calls ZSTD_compress() and RestoreBlockImage() > calls ZSTD_decompress(). Both create and destroy a context per call. At > the default level ZSTD_estimateCCtxSize() reports 1.3MB for the > contex. So zstd pays an allocation per full-page image on the > way in, and one per image again on the way out. The patch creates each > on first use and keeps it, the compressor for the life of the backend, > the decompressor in XLogReaderState. > > The reading side is where this shows, because one startup process > replays every image. Redoing 200MB of page images: > > zstd, master 4.24s zstd, patched 2.33s > lz4, master 1.59s lz4, patched 1.59s > > A sequential scan that sets hint bits on a freshly checkpointed table > with wal_log_hints on runs about 26% faster. > > This also bears on the "WAL compression setting after PostgreSQL LZ4 > default change" thread, where the order zstd -> lz4 -> pglz is proposed > for what "on" should mean. On master, replaying zstd-compressed images > costs 2.7x what lz4 costs, even though zstd wrote 38% less WAL; with the > context kept, that gap falls to 1.5x. > > Cons: a backend that used zstd once holds the 1.3MB until it exits, and > a reader holds a decompression context, which is far smaller. > > ---- 0002: whole-record compression alongside FPI compression ---- > > Unchanged in design from v7, plus the documentation that was missing. > > When a record is larger than wal_compression_threshold it is compressed > as a single unit rather than each full-page image separately, which wins > whenever the images in one record share content. > > On CREATE INDEX over 10M random doubles, with the generator seeded so > the runs are comparable: > > zstd 144.8MB -> 116.9MB (-19%) > lz4 189.5MB -> 161.0MB (-15%) > > Pros: it reaches redundancy that per-FPI compression cannot see, records > below the threshold are untouched, and every record still decodes on its > own, so nothing about how WAL is read changes. Setting the threshold > above the largest possible record restores today's behaviour. > > Cons, both unchanged from v7: > > 1. Memory. A backend with compression enabled holds two 274300-byte > buffers where it used to hold one array of about the same total size: > the staging buffer reproduces the old allocation, the output buffer > is new. The compressed length has to be known before WAL space is > reserved, so the output has to be materialized somewhere. > > 2. A low threshold can make WAL bigger, because whole-record compression > displaces per-FPI compression for the records it takes and adds > header to each. > > > ---- 0003: compress records against earlier records (WIP) ---- > > Still WIP, but no longer blocked on the reader problem. > > The motivation is unchanged: half the WAL a pgbench run produces is out > of reach for per-record compression, because the median record is far > too small to compress on its own but compresses well against the records > before it. > > What is new in v8 is how a reader starts in the middle of WAL. > > Every stream starts over at fixed 4MB boundaries. The writer enforces > that rather than hoping for it: a record that would continue a stream > past the next boundary is refused its reserved position and built again > against a stream that starts over. The check is one comparison against > CurrBytePos under the spinlock that reservation already holds; the > boundary is converted to a byte position outside the lock. It fires > about once per stream per 4MB, and I could not measure its cost. > > A reader that wants to start at some LSN then rewinds to the boundary > below it and reads forward, which rebuilds the decompressors, and stops > short of the record it was asked for: feeding a record to its > decompressor twice would leave it in a state its successors were not > compressed against. That is XLogBeginReadStreamed(), and pg_waldump, > pg_walinspect, logical decoding, walsummarizer and pg_rewind all use it. > A record whose stream has not been seen to start over refuses to > decompress rather than decoding whatever the bytes happen to mean. > > Why a fixed distance and not the WAL segment: how far a reader rewinds > should not change when a cluster is initialised with a different segment > size, and 64MB segments are not unusual. 4MB divides both 16MB and 64MB > segments, so a boundary is also always a page start. Compression is > insensitive to the value - pgbench emits the same WAL per transaction > to within 2% anywhere between 1MB and 64MB - so it is chosen for the > readers. > > Numbers below are from a 16-core machine, pgbench scale 100, fsync off > so that this measures the feature and not the disk. WAL bytes per > transaction and throughput, with the stream count matched to the client > count: > > clients streams off streams = clients > 1 1720 / 2456 1086 / 2340 -37% WAL, -5% tps > 8 701 / 18432 530 / 17326 -24% WAL, -6% tps > 32 544 / 43729 433 / 44111 -20% WAL, no cost > 64 542 / 46173 448 / 39499 -17% WAL, -12% tps > > The 32- and 64-client rows answer the question I could not answer in v7: > the stream lease, which is held across compression and insertion, does > not show up as contention when the insert path is already busy. At 32 > clients with 32 streams there is no measurable cost at all. > > What does show up is that the stream count has to track concurrency. > Eight streams buy 4% at 32 clients and nothing at 64: > > 64 clients, streams 0 / 8 / 64 -> 542 / 534 / 448 bytes per txn > > Building the third patch with wal_compression_streams = 0 reproduces the > second patch's numbers, so the cost is in using the feature, not in > carrying it. > > Two other shapes, WAL volume only: > > wide UPDATE of 500k rows 223.9MB -> 154.6MB (-31%) > COPY of 3M rows 65.4MB -> 68.6MB (+5%) > > The COPY case is the honest counterexample: those records are large and > already compress well on their own, so the stream adds header and buys > nothing. I do not think that argues against the feature, but it does > argue that turning it on should stay a choice. > > Memory, peak RSS with 64 clients writing and then a full pg_waldump over > what they wrote: > > streams backend peak pg_waldump peak > 0 146.6 MB 3.4 MB > 8 150.5 MB 9.0 MB > 64 150.5 MB 51.5 MB > > The writing side costs about 4MB per backend and does not grow with the > stream count, because a backend keeps one compressor rather than one per > stream. The reading side costs about 0.75MB per stream, and every > reader pays it: the startup process, a walsender doing logical decoding, > pg_waldump. > > Costs, as I see them: > > 1. Throughput, when streams are pushed as high as the client count: 12% > at 64 clients on 16 cores. At and below one stream per core I could > not measure a cost. > > 2. Memory. 0.75MB per stream for every reader, as above. > > 3. WAL retention. A replication slot has to keep the WAL back to the > reset boundary below what it needs itself, so up to 4MB more. > > The same rewind costs reading, not just retention: a reader that > starts in the middle re-reads up to 4MB to rebuild the decompressors, > and it does so whether or not the WAL it is about to read holds any > streams at all. For readers that start once that is nothing; for > walsummarizer, which starts afresh per summary file, it is up to 4MB > per 16MB summarized. Making the rewind happen only when a record > actually turns out to need it is the obvious answer and I have not > got it working yet. > > 4. Records that someone reads by LSN without replaying what precedes > them have to stay out of the scheme: the checkpoint records including > XLOG_CHECKPOINT_REDO, XLOG_END_OF_RECOVERY, XLOG_SWITCH and PREPARE. > > I could not measure the fsync=on case usefully. The disk I have caps at > 82MB/s and repeats of one configuration differed by a factor of two, so > I have no throughput claim there; WAL volume did reproduce, 15-25% > lower with streams, in line with the numbers above. > > What I would still like opinions on: whether refusing a reserved > position is an acceptable thing for an insertion path to do, and how to > resolve the tension the numbers above show: the ratio wants roughly one > stream per writing backend, while each stream costs every reader > 0.75MB. > wal_compression_streams currently caps at 64, which is already generous > for a reader and far short of the backend count on a busy server. > > WDYT? The patches cannot be applied to the current tree because they conflict with commits 931c9701f25, 18992dc9d98, and b614de4876b. Below are some initial review comments. v8-0002 1. + report_invalid_record(state, + "could not decompress record at %X/%08X compressed with %s not supported by build", + LSN_FORMAT_ARGS((XLogRecPtr) recptr), "lz4"); + return NULL; The casting of XLogRecPtr is unnecessary. 2. + report_invalid_record(state, + "could not decompress record at %X/%08X compressed with %s not supported by build", + LSN_FORMAT_ARGS((XLogRecPtr) recptr), "zstd"); + return NULL; Same as above. 3. + report_invalid_record(state, + "could not decompress record at %X/%08X compressed with unknown method", + LSN_FORMAT_ARGS((XLogRecPtr) recptr)); + return NULL; Same as above. v8-0003 1. + state->stream_dctx = palloc0(sizeof(void *) * XLR_MAX_STREAMS); + state->stream_ready = palloc0(sizeof(bool) * XLR_MAX_STREAMS); I'd prefer use the palloc0_array() macro. > > > Best regards, Andrey Borodin. -- Regards, Japin Li ChengDu WenWu Information Technology Co., Ltd.
