On Wed, Sep 23, 2026 at 1:09 PM Nathan Bossart <[email protected]> wrote: > > On Wed, Sep 23, 2026 at 09:41:36AM -0500, Nathan Bossart wrote: > > On Wed, Sep 23, 2026 at 10:10:09AM -0400, Sehrope Sarkuni wrote: > >> lpad() and rpad() pad one character at a time, calling > >> pg_mblen_range() and memcpy() once per padding char. When the padding > >> string is a single byte, e.g., lpad(x, n, '0') or rpad(x, n, ' '), > >> the padding is that byte repeated, so the attached patch fills it with > >> one memset(). > > > > I wonder if we could expand these gains by using SIMD whenever the vector > > length is divisible by the padding string length. My hunch is that's where > > a lot of the memset() gains come from. > > Actually, I think we can expand this to any padding string length by > copying the padding string once, and then copying from the beginning of the > padding to the end repeatedly so that we write double the padding each > time. This is a bit like what commit c60e520 added for pglz_decompress(). > I've attached some proof-of-concept grade patches. This doesn't quite > match the performance of your 1-byte fast-path, but it's pretty close and > applies to many more cases.
Ah! just saw this after I hit send. I think my version does the rest of what you're describing! Exact conclusion reached from testing it too. Regards, -- Sehrope Sarkuni Founder & CEO | JackDB, Inc. | https://www.jackdb.com/
