On Wed, Sep 23, 2026 at 1:09 PM Nathan Bossart <[email protected]> wrote:
>
> On Wed, Sep 23, 2026 at 09:41:36AM -0500, Nathan Bossart wrote:
> > On Wed, Sep 23, 2026 at 10:10:09AM -0400, Sehrope Sarkuni wrote:
> >> lpad() and rpad() pad one character at a time, calling
> >> pg_mblen_range() and memcpy() once per padding char.  When the padding
> >> string is a single byte, e.g., lpad(x, n, '0') or rpad(x, n, ' '),
> >> the padding is that byte repeated, so the attached patch fills it with
> >> one memset().
> >
> > I wonder if we could expand these gains by using SIMD whenever the vector
> > length is divisible by the padding string length.  My hunch is that's where
> > a lot of the memset() gains come from.
>
> Actually, I think we can expand this to any padding string length by
> copying the padding string once, and then copying from the beginning of the
> padding to the end repeatedly so that we write double the padding each
> time.  This is a bit like what commit c60e520 added for pglz_decompress().
> I've attached some proof-of-concept grade patches.  This doesn't quite
> match the performance of your 1-byte fast-path, but it's pretty close and
> applies to many more cases.

Ah! just saw this after I hit send.

I think my version does the rest of what you're describing!

Exact conclusion reached from testing it too.

Regards,
-- Sehrope Sarkuni
Founder & CEO | JackDB, Inc. | https://www.jackdb.com/


Reply via email to