On Fri, Aug 14, 2026 at 6:08 PM Masahiko Sawada <[email protected]> wrote: > > On Fri, Aug 14, 2026 at 4:33 PM Masahiko Sawada <[email protected]> wrote: > > > > On Thu, Aug 6, 2026 at 8:07 PM Chao Li <[email protected]> wrote: > > > > > > > > > > > > > On Aug 7, 2026, at 00:09, Masahiko Sawada <[email protected]> wrote: > > > > > > > > On Wed, Aug 5, 2026 at 8:40 PM Chao Li <[email protected]> wrote: > > > >> > > > >> > > > >> > > > >>> On Aug 6, 2026, at 08:19, Masahiko Sawada <[email protected]> > > > >>> wrote: > > > >>> > > > >>> On Tue, Jun 30, 2026 at 11:03 AM Haibo Yan <[email protected]> > > > >>> wrote: > > > >>>> > > > >>>> On Tue, Jun 30, 2026 at 10:53 AM Masahiko Sawada > > > >>>> <[email protected]> wrote: > > > >>>>> > > > >>>>> On Mon, Jun 29, 2026 at 5:53 PM Haibo Yan <[email protected]> > > > >>>>> wrote: > > > >>>>>> > > > >>>>>> On Mon, Jun 29, 2026 at 2:55 PM Masahiko Sawada > > > >>>>>> <[email protected]> wrote: > > > >>>>>>> > > > >>>>>>> On Sun, Jun 28, 2026 at 7:20 PM Haibo Yan <[email protected]> > > > >>>>>>> wrote: > > > >>>>>>>> > > > >>>>>>>> On Thu, Jun 25, 2026 at 3:16 PM Masahiko Sawada > > > >>>>>>>> <[email protected]> wrote: > > > >>>>>>>>> > > > >>>>>>>>> On Thu, Jun 25, 2026 at 2:31 PM Haibo Yan > > > >>>>>>>>> <[email protected]> wrote: > > > >>>>>>>>>> > > > >>>>>>>>>> > > > >>>>>>>>>> > > > >>>>>>>>>> On Thu, Jun 25, 2026 at 11:28 AM Masahiko Sawada > > > >>>>>>>>>> <[email protected]> wrote: > > > >>>>>>>>>>> > > > >>>>>>>>>>> Hi all, > > > >>>>>>>>>>> > > > >>>>>>>>>>> I'd like to propose the $subject. > > > >>>>>>>>>>> > > > >>>>>>>>>>> Since commit ec8719ccbfcd made hex_decode_safe() SIMD-aware, > > > >>>>>>>>>>> decoding > > > >>>>>>>>>>> a run of hex digits is now fast. The attached patch reuses > > > >>>>>>>>>>> hex_decode_safe() in the UUID input function to speed up > > > >>>>>>>>>>> parsing. > > > >>>>>>>>>>> > > > >>>>>>>>>>> We accept several textual forms of a UUID[1]. The fast path > > > >>>>>>>>>>> handles > > > >>>>>>>>>>> the common ones: 32 hex digits, the canonical 8x-4x-4x-4x-12x > > > >>>>>>>>>>> form > > > >>>>>>>>>>> (where "nx" means n hex digits), and either of those wrapped > > > >>>>>>>>>>> in > > > >>>>>>>>>>> braces. Otherwise, it falls back to the ordinary scalar UUID > > > >>>>>>>>>>> parse. > > > >>>>>>>>>>> > > > >>>>>>>>>>> I've benchmarked the parse speed using the following query: > > > >>>>>>>>>>> > > > >>>>>>>>>>> CREATE TEMP TABLE u AS SELECT gen_random_uuid()::text AS t > > > >>>>>>>>>>> FROM > > > >>>>>>>>>>> generate_series(1, 1000000); > > > >>>>>>>>>>> EXPLAIN (ANALYZE, TIMING OFF) SELECT t::uuid FROM u; > > > >>>>>>>>>>> > > > >>>>>>>>>>> I compared the execution time of the second query, which > > > >>>>>>>>>>> measures > > > >>>>>>>>>>> uuid_in() alone, with/without SIMD optimization. Here are > > > >>>>>>>>>>> results (the > > > >>>>>>>>>>> median of 5 runs): > > > >>>>>>>>>>> > > > >>>>>>>>>>> HEAD: 208.879 ms > > > >>>>>>>>>>> Patched: 40.983 ms > > > >>>>>>>>>>> > > > >>>>>>>>>>> The improvements look promising to me. But in a realistic > > > >>>>>>>>>>> pipeline the > > > >>>>>>>>>>> parse is a small fraction of the work, so end-to-end gains > > > >>>>>>>>>>> could be > > > >>>>>>>>>>> much smaller. > > > >>>>>>>>>>> > > > >>>>>>>>>>> Feedback is very welcome. > > > >>>>>>>>>>> > > > >>>>>>>>>> I may be missing something, but I wonder whether the fast path > > > >>>>>>>>>> is relying on > > > >>>>>>>>>> slightly different input semantics from the existing UUID > > > >>>>>>>>>> parser. > > > >>>>>>>>>> > > > >>>>>>>>>> In particular, hex_decode_safe() is not a strict “32 hex > > > >>>>>>>>>> characters only” > > > >>>>>>>>>> decoder. It skips whitespace, which is fine for its existing > > > >>>>>>>>>> callers, but I > > > >>>>>>>>>> don’t think UUID input should treat whitespace inside the UUID > > > >>>>>>>>>> body as > > > >>>>>>>>>> ignorable. > > > >>>>>>>>> > > > >>>>>>>>> Good catch! hex_decode_safe() skips whitespaces so the patch > > > >>>>>>>>> accepts > > > >>>>>>>>> the following UUID value, which is bad: > > > >>>>>>>>> > > > >>>>>>>>> select '019f00b5-7f8a-722f-b707-59f0ed25cd '::uuid; > > > >>>>>>>>> uuid > > > >>>>>>>>> -------------------------------------- > > > >>>>>>>>> 019f00b5-7f8a-722f-b707-59f0ed25cd00 > > > >>>>>>>>> (1 row) > > > >>>>>>>>> > > > >>>>>>>>>> Also, since hex_decode_safe() returns void, the UUID fast path > > > >>>>>>>>>> cannot verify that exactly UUID_LEN bytes were produced. > > > >>>>>>>>> > > > >>>>>>>>> IIUC hex_decode_safe() does return the output length in bytes. > > > >>>>>>>>> So I > > > >>>>>>>>> think we can fallback to the scalar UUID parser if > > > >>>>>>>>> esctx.error_occurred is true or if the returned value is not 16. > > > >>>>>>>>> > > > >>>>>>>> > > > >>>>>>>> You’re right, I misread that part. Checking both > > > >>>>>>>> esctx.error_occurred and > > > >>>>>>>> the returned length sounds good to me. > > > >>>>>>>> > > > >>>>>>>>>> > > > >>>>>>>>>> So I think it would be safer either to pre-validate that the > > > >>>>>>>>>> 32 source > > > >>>>>>>>>> characters are all hex digits before calling > > > >>>>>>>>>> hex_decode_safe(), or to use a > > > >>>>>>>>>> UUID-specific strict hex decoder for this path. After that, a > > > >>>>>>>>>> comment > > > >>>>>>>>>> explaining why hex_decode_safe() is safe here would make the > > > >>>>>>>>>> invariant much > > > >>>>>>>>>> clearer. > > > >>>>>>>>> > > > >>>>>>>>> IIUC hex_decode_simd_helper() accepts only hex digits so we > > > >>>>>>>>> could > > > >>>>>>>>> re-use it for UUID parsing. Let me check if the above idea of > > > >>>>>>>>> using > > > >>>>>>>>> the return value works for us first. > > > >>>>>>>>> > > > >>>>>>>> > > > >>>>>>>> That sounds reasonable. My main concern was to keep the fast > > > >>>>>>>> path’s accepted > > > >>>>>>>> input set identical to the scalar UUID parser. Falling back > > > >>>>>>>> when the decoded > > > >>>>>>>> length is not UUID_LEN, together with regression tests for > > > >>>>>>>> whitespace cases, > > > >>>>>>>> should address that. > > > >>>>>>>> > > > >>>>>>>>>> > > > >>>>>>>>>> Could you also add a few regression tests for invalid inputs > > > >>>>>>>>>> that contain > > > >>>>>>>>>> whitespace inside otherwise fast-path-looking UUID strings? > > > >>>>>>>>>> For example: > > > >>>>>>>>>> > > > >>>>>>>>>> --------------------------------------------------------------- > > > >>>>>>>>>> > > > >>>>>>>>>> SELECT 'a0eebc99 9c0b4ef8bb6d6bb9bd380a11'::uuid; > > > >>>>>>>>>> SELECT 'a0eebc999c0b4ef8bb6d6bb9bd380a1 '::uuid; > > > >>>>>>>>>> SELECT '{a0eebc999c0b4ef8bb6d6bb9bd380a1 }'::uuid; > > > >>>>>>>>>> SELECT 'a0eebc99-9c0b-4ef8-bb6d-6bb9bd380a1 '::uuid; > > > >>>>>>>>>> --------------------------------------------------------------- > > > >>>>>>>>>> > > > >>>>>>>>>> These should continue to be rejected in the same way as the > > > >>>>>>>>>> scalar parser. > > > >>>>>>>>>> Regards, > > > >>>>>>>>> > > > >>>>>>>>> Agreed. > > > >>>>>>>>> > > > >>>>>>> > > > >>>>>>> I've attached the updated patch. > > > >>>>>>> > > > >>>>>>> Regards, > > > >>>>>>> > > > >>>>>>> -- > > > >>>>>>> Masahiko Sawada > > > >>>>>>> Amazon Web Services: https://aws.amazon.com > > > >>>>>> > > > >>>>>> I noticed a few typos in the comments: > > > >>>>>> > > > >>>>>> src/backend/utils/adt/uuid.c > > > >>>>>> line 56: “scalar implmentation” -> “scalar implementation” > > > >>>>>> line 109: “swalled” -> “swallowed” > > > >>>>>> line 110: “kepping” -> “keeping” > > > >>>>>> line 118: “grammer” -> “grammar” > > > >>>>>> line 119: “whitespaces” -> “whitespace” > > > >>>>>> > > > >>>>>> Could you fix them ? > > > >>>>> > > > >>>>> Oops, I fixed them and rechecked other places. > > > >>>>> > > > >>>>> I've attached the updated patch. > > > >>>>> > > > >>>>> Regards, > > > >>>>> > > > >>>>> -- > > > >>>>> Masahiko Sawada > > > >>>>> Amazon Web Services: https://aws.amazon.com > > > >>>> > > > >>>> The code looks good to me now. I only noticed one small typo in the > > > >>>> commit trailer: Reviwed-by should be Reviewed-by. > > > >>>> > > > >>>> Otherwise, it looks good. Thank you for fixing these issues. > > > >>>> > > > >>> > > > >>> After spending more time on this patch, I find out two things: > > > >>> > > > >>> 1. USE_NO_SIMD doesn't work in uuid.c without including port/simd.h. > > > >>> But including port/simd.h seems wrong as it doesn't use any SIMD > > > >>> support functions. > > > >>> > > > >>> 2. hex_decode_safe() is faster than the current UUID parse > > > >>> (isxdigit()+strtoul() approach) even without SIMD. I've created a > > > >>> small benchmark test tool (attached as 0002 patch, not intended to be > > > >>> pushed into the core), and measures UUID parsing performance of three > > > >>> approaches: 'scalar' is the current string_to_uuid() that uses > > > >>> isxdigit()+strtoul()), 'simd' uses hex_decode_safe() with SIMD, and > > > >>> 'nosimd' uses hex_decode_safe() without SIMD, with different shapes of > > > >>> UUIDs. Here are results: > > > >>> > > > >>> =# select path, shape, n_inputs, best_ms::numeric(10,3) from > > > >>> uuid_parse_bench(100000, 5); > > > >>> path | shape | n_inputs | best_ms > > > >>> --------+------------------+----------+--------- > > > >>> scalar | canonical | 100000 | 22.661 > > > >>> simd | canonical | 100000 | 1.400 > > > >>> nosimd | canonical | 100000 | 1.652 > > > >>> scalar | bare32 | 100000 | 15.932 > > > >>> simd | bare32 | 100000 | 0.471 > > > >>> nosimd | bare32 | 100000 | 1.110 > > > >>> scalar | braced_canonical | 100000 | 17.330 > > > >>> simd | braced_canonical | 100000 | 1.088 > > > >>> nosimd | braced_canonical | 100000 | 1.314 > > > >>> scalar | braced_bare32 | 100000 | 15.942 > > > >>> simd | braced_bare32 | 100000 | 0.488 > > > >>> nosimd | braced_bare32 | 100000 | 1.141 > > > >>> scalar | dashed4 | 100000 | 16.185 > > > >>> simd | dashed4 | 100000 | 16.493 > > > >>> nosimd | dashed4 | 100000 | 16.403 > > > >>> scalar | invalid_hex | 100000 | 0.199 > > > >>> simd | invalid_hex | 100000 | 1.150 > > > >>> nosimd | invalid_hex | 100000 | 0.385 > > > >>> (18 rows) > > > >>> > > > >>> Each of shape means: > > > >>> - 'canonical': 8x-4x-4x-4x-12x, what uuid_out() emits > > > >>> - 'bare32': 32 contiguous hex digits > > > >>> - 'braced_canonical': {8x-4x-4x-4x-12x} > > > >>> - 'braced_bare32': {32 hex digits} > > > >>> - 'dashed4': dash after every group of 4 > > > >>> - 'invalid_hdx': canonical but with a invalid digit > > > >>> > > > >>> 'nosimd' is 10x~ faster than 'scalar' in most cases. All paths are > > > >>> mostly the same in 'dashed4' and 'invalid_hex' cases because 'simd' > > > >>> and 'nosimd' fall back to the 'scalar' case. According to these > > > >>> results, my conclusion is that we can use hex_decode_safe() for > > > >>> canonical forms and 32 contiguous hex forms anyway, and let > > > >>> hex_decode_safe() choose whether to use SIMD. We would win in either > > > >>> case. We still use the current scalar approach for uncommon UUID forms > > > >>> and error reporting purposes. > > > >>> > > > >>> Regards, > > > >>> > > > >>> -- > > > >>> Masahiko Sawada > > > >>> Amazon Web Services: https://aws.amazon.com > > > >>> <v4-0001-Optimize-UUID-parse-using-SIMD.patch><v4-0002-uuid_parse_bench-module.patch> > > > >> > > > >> A few comments on v4. > > > >> > > > >> 1 - 0001 > > > >> ``` > > > >> +static void > > > >> +string_to_uuid(const char *source, pg_uuid_t *uuid, Node *escontext) > > > >> +{ > > > >> + const char *body = source; > > > >> + size_t len = strlen(source); > > > >> ``` > > > >> > > > >> I think it would be better to avoid strlen(). The old code processes > > > >> at most UUID_LEN (16) byte pairs, so it does not need to scan > > > >> arbitrarily far on malformed input. So, maybe we could use something > > > >> like strnlen(source, 39) instead. > > > > > > > > While strnlen(source, 39) works there, 39 is a magic number and it's > > > > tied to the current format check logic. What is the benefit of using > > > > strnlen(source, 39) instead? I'm not sure it warrants having the magic > > > > number. > > > > > > It doesn't have to be exactly 39; 1024 (long enough) would also work, or > > > perhaps something based on UUID_LEN, such as UUID_LEN * 3. I think the > > > main point is to avoid unbounded scanning on malformed input. > > > > > > The old code did not have this issue because it only examined as much > > > input as needed based on UUID_LEN. The new fast path starts to use > > > strlen(), so this would be a new risk introduced by the optimization. > > > > I don't think the scan can be really unbounded. string_to_uuid() > > receives a cstring, so by the time it is called the caller has already > > walked or copied the whole string to produce it. So unless the > > unbounded scan can be reached in some path I have overlooked, I'd > > prefer to keep strlen() here. Happy to change it if you still think it > > is worth it. > > After more thoughts, while I still don't think the scan can be > unbounded, using strlen() would add an extra scan just to determine we > use hex_decode_safe(). I'll change it to use strnlen() instead.
I've updated the patch accordingly. Please review it. Regards, -- Masahiko Sawada Amazon Web Services: https://aws.amazon.com
From 52e9030780fadcaf905062ceefecb4f87737ab9d Mon Sep 17 00:00:00 2001 From: Masahiko Sawada <[email protected]> Date: Thu, 25 Jun 2026 10:03:44 -0700 Subject: [PATCH v5 1/2] Optimize UUID parse using SIMD. Previously, string_to_uuid() decoded one byte at a time, calling isxdigit() twice and strtoul() once for every pair of hexadecimal digits. That loop dominated the cost of uuid_in(). This commit adds a fast path for the two common shapes: a bare string of 32 hexadecimal digits, and the canonical 8x-4x-4x-4x-12x form (where "nx" means n hexadecimal digits), each optionally wrapped in braces. Both are compacted into 32 contiguous hexadecimal digits and decoded with hex_decode_safe(). Any other shape, or any decoding error, is handed off to the original scalar parser, now string_to_uuid_scalar(), so the accepted grammar and the error messages are unchanged. hex_decode_safe() silently skips whitespace while the UUID grammar does not, so a decode can succeed and still write fewer than UUID_LEN bytes. The fast path therefore treats a short result as a failure, just like an error, and lets the scalar parser reject the input and report the syntax error. The fast path is deliberately not conditional on SIMD support. hex_decode_safe() selects a vectorized or scalar implementation itself, and even its scalar implementation is an order of magnitude faster than decoding a byte at a time, so gating this on USE_NO_SIMD would only penalize platforms that have neither SSE2 nor NEON. Reviewed-by: Bharath Rupireddy <[email protected]> Reviewed-by: Haibo Yan <[email protected]> Discussion: https://postgr.es/m/cad21aocqer4uqu77q_yomnnzj7aveio5qzt+4hnzpm4wm-e...@mail.gmail.com --- src/backend/utils/adt/uuid.c | 103 +++++++++++++++++++++++++++-- src/test/regress/expected/uuid.out | 75 +++++++++++++++++++++ src/test/regress/sql/uuid.sql | 25 +++++++ 3 files changed, 198 insertions(+), 5 deletions(-) diff --git a/src/backend/utils/adt/uuid.c b/src/backend/utils/adt/uuid.c index 28e18940a9d..9a3b6715828 100644 --- a/src/backend/utils/adt/uuid.c +++ b/src/backend/utils/adt/uuid.c @@ -19,7 +19,9 @@ #include "common/hashfn.h" #include "lib/hyperloglog.h" #include "libpq/pqformat.h" +#include "nodes/miscnodes.h" #include "port/pg_bswap.h" +#include "utils/builtins.h" #include "utils/fmgrprotos.h" #include "utils/guc.h" #include "utils/skipsupport.h" @@ -139,13 +141,13 @@ uuid_out(PG_FUNCTION_ARGS) } /* - * We allow UUIDs as a series of 32 hexadecimal digits with an optional dash - * after each group of 4 hexadecimal digits, and optionally surrounded by {}. - * (The canonical format 8x-4x-4x-4x-12x, where "nx" means n hexadecimal - * digits, is the only one used for output.) + * Reference implementation of the UUID grammar, parsing one character at a + * time. string_to_uuid() recognizes the common shapes more cheaply and + * defers to this function for everything else, so this is also the only + * place that reports a syntax error. */ static void -string_to_uuid(const char *source, pg_uuid_t *uuid, Node *escontext) +string_to_uuid_scalar(const char *source, pg_uuid_t *uuid, Node *escontext) { const char *src = source; bool braces = false; @@ -194,6 +196,97 @@ syntax_error: "uuid", source))); } +/* + * We allow UUIDs as a series of 32 hexadecimal digits with an optional dash + * after each group of 4 hexadecimal digits, and optionally surrounded by {}. + * (The canonical format 8x-4x-4x-4x-12x, where "nx" means n hexadecimal + * digits, is the only one used for output.) + * + * The two common shapes -- a bare string of 32 hexadecimal digits and the + * canonical form, each optionally wrapped in braces -- are compacted into 32 + * contiguous hex digits and decoded with hex_decode_safe(), which is much + * faster than the character-at-a-time loop. Any other shape, or any decoding + * error, is handed off to string_to_uuid_scalar() so that the accepted + * grammar and the error messages are unchanged. + * + * Note that this fast path is not conditional on SIMD support: + * hex_decode_safe() picks a vectorized or scalar implementation itself, and + * even its scalar implementation is far faster than string_to_uuid_scalar(). + */ +static void +string_to_uuid(const char *source, pg_uuid_t *uuid, Node *escontext) +{ + const char *body = source; + const char *hexsrc = NULL; + char hexbuf[32]; + uint64 written; + size_t len; + ErrorSaveContext esctx = {T_ErrorSaveContext}; + + /* + * Measure the input only far enough to classify its shape. The bound + * must exceed the longest shape handled here, the braced canonical form + * at 38 characters: strnlen() returns the bound for anything at least + * that long, so stopping at an accepted length would accept a longer + * string that merely starts with a valid UUID. + */ + len = strnlen(source, 64); + + /* Strip one optional surrounding brace pair */ + if (len >= 2 && source[0] == '{' && source[len - 1] == '}') + { + body = source + 1; + len -= 2; + } + + if (len == 32) + { + /* + * Body is already 32 contiguous hex digits -- decode straight from + * the input. hex_decode_safe() reads exactly body[0..31], so it never + * touches the trailing NULL or '}'. + */ + hexsrc = body; + } + else if (len == 36 && body[8] == '-' && body[13] == '-' && + body[18] == '-' && body[23] == '-') + { + /* + * Canonical 8x-4x-4x-4x-12x form; compact them into hexbuf with + * fixed-offset copies, dropping the dashes. + */ + memcpy(&hexbuf[0], &body[0], 8); + memcpy(&hexbuf[8], &body[9], 4); + memcpy(&hexbuf[12], &body[14], 4); + memcpy(&hexbuf[16], &body[19], 4); + memcpy(&hexbuf[20], &body[24], 12); + hexsrc = hexbuf; + } + + if (hexsrc == NULL) + { + /* Uncommon shape; let the general parse handle it */ + string_to_uuid_scalar(source, uuid, escontext); + return; + } + + /* + * Decode the UUID hex data using our hex decoder that is SIMD-aware. We + * give it a private error context so that a decode failure is swallowed + * here and reported by the scalar path instead, keeping the error message + * identical. + */ + written = hex_decode_safe(hexsrc, 32, (char *) uuid->data, (Node *) &esctx); + + /* + * Fall back to the scalar path on any error. We must also reject a short + * result: hex_decode_safe() skips whitespace, so it can succeed yet write + * fewer than UUID_LEN bytes, whereas the UUID grammar forbids whitespace. + */ + if (esctx.error_occurred || written != UUID_LEN) + string_to_uuid_scalar(source, uuid, escontext); +} + Datum uuid_recv(PG_FUNCTION_ARGS) { diff --git a/src/test/regress/expected/uuid.out b/src/test/regress/expected/uuid.out index d542eb14b26..6247bae8574 100644 --- a/src/test/regress/expected/uuid.out +++ b/src/test/regress/expected/uuid.out @@ -375,5 +375,80 @@ SELECT v = v::bytea::uuid as matched FROM gen_random_uuid() v; t (1 row) +-- Test UUID shapes that the parser uses the SIMD path. +SELECT '5b35380a-7143-4912-9b55-f322699c6770'::uuid; + uuid +-------------------------------------- + 5b35380a-7143-4912-9b55-f322699c6770 +(1 row) + +SELECT '{5b35380a-7143-4912-9b55-f322699c6770}'::uuid; + uuid +-------------------------------------- + 5b35380a-7143-4912-9b55-f322699c6770 +(1 row) + +SELECT '5b35380a714349129b55f322699c6770'::uuid; + uuid +-------------------------------------- + 5b35380a-7143-4912-9b55-f322699c6770 +(1 row) + +SELECT '{5b35380a714349129b55f322699c6770}'::uuid; + uuid +-------------------------------------- + 5b35380a-7143-4912-9b55-f322699c6770 +(1 row) + +-- Test if the UUID parser using SIMD optimization correctly rejects invalid UUID +-- string format. +SELECT '5b35380a714349129b55f32 99c6770'::uuid; +ERROR: invalid input syntax for type uuid: "5b35380a714349129b55f32 99c6770" +LINE 1: SELECT '5b35380a714349129b55f32 99c6770'::uuid; + ^ +SELECT '5b35380a-7143-4912-9b55-f322699c67 '::uuid; +ERROR: invalid input syntax for type uuid: "5b35380a-7143-4912-9b55-f322699c67 " +LINE 1: SELECT '5b35380a-7143-4912-9b55-f322699c67 '::uuid; + ^ +SELECT ' 35380a-7143-4912-9b55-f322699c6770'::uuid; +ERROR: invalid input syntax for type uuid: " 35380a-7143-4912-9b55-f322699c6770" +LINE 1: SELECT ' 35380a-7143-4912-9b55-f322699c6770'::uuid; + ^ +SELECT 'AZ35380a-7143-4912-9b55-f322699c6770'::uuid; +ERROR: invalid input syntax for type uuid: "AZ35380a-7143-4912-9b55-f322699c6770" +LINE 1: SELECT 'AZ35380a-7143-4912-9b55-f322699c6770'::uuid; + ^ +SELECT '{AZ35380a-7143-4912-9b55-f322699c6770}'::uuid; +ERROR: invalid input syntax for type uuid: "{AZ35380a-7143-4912-9b55-f322699c6770}" +LINE 1: SELECT '{AZ35380a-7143-4912-9b55-f322699c6770}'::uuid; + ^ +SELECT '{AZ35380a714349129b55f322699c6770}'::uuid; +ERROR: invalid input syntax for type uuid: "{AZ35380a714349129b55f322699c6770}" +LINE 1: SELECT '{AZ35380a714349129b55f322699c6770}'::uuid; + ^ +SELECT '{AZ35380a714349129b55f322699c67 }'::uuid; +ERROR: invalid input syntax for type uuid: "{AZ35380a714349129b55f322699c67 }" +LINE 1: SELECT '{AZ35380a714349129b55f322699c67 }'::uuid; + ^ +-- The parser only measures the input far enough to classify its shape. If it +-- stopped measuring at one of the accepted lengths, a longer string that +-- merely starts with a valid UUID would look like that UUID and be accepted +-- with the rest silently ignored, so check that trailing data is rejected. +SELECT '5b35380a714349129b55f322699c6770TRAILING'::uuid; +ERROR: invalid input syntax for type uuid: "5b35380a714349129b55f322699c6770TRAILING" +LINE 1: SELECT '5b35380a714349129b55f322699c6770TRAILING'::uuid; + ^ +SELECT '{5b35380a714349129b55f322699c6770}TRAILING'::uuid; +ERROR: invalid input syntax for type uuid: "{5b35380a714349129b55f322699c6770}TRAILING" +LINE 1: SELECT '{5b35380a714349129b55f322699c6770}TRAILING'::uuid; + ^ +SELECT '5b35380a-7143-4912-9b55-f322699c6770TRAILING'::uuid; +ERROR: invalid input syntax for type uuid: "5b35380a-7143-4912-9b55-f322699c6770TRAILING" +LINE 1: SELECT '5b35380a-7143-4912-9b55-f322699c6770TRAILING'::uuid; + ^ +SELECT '{5b35380a-7143-4912-9b55-f322699c6770}TRAILING'::uuid; +ERROR: invalid input syntax for type uuid: "{5b35380a-7143-4912-9b55-f322699c6770}TRAILING" +LINE 1: SELECT '{5b35380a-7143-4912-9b55-f322699c6770}TRAILING'::uui... + ^ -- clean up DROP TABLE guid1, guid2, guid3 CASCADE; diff --git a/src/test/regress/sql/uuid.sql b/src/test/regress/sql/uuid.sql index 54f0d8f8255..d13e8c21921 100644 --- a/src/test/regress/sql/uuid.sql +++ b/src/test/regress/sql/uuid.sql @@ -178,5 +178,30 @@ SELECT '\x019a2f859ced7225b99d9c55044a2563'::bytea::uuid; SELECT '\x1234567890abcdef'::bytea::uuid; -- error SELECT v = v::bytea::uuid as matched FROM gen_random_uuid() v; +-- Test UUID shapes that the parser uses the SIMD path. +SELECT '5b35380a-7143-4912-9b55-f322699c6770'::uuid; +SELECT '{5b35380a-7143-4912-9b55-f322699c6770}'::uuid; +SELECT '5b35380a714349129b55f322699c6770'::uuid; +SELECT '{5b35380a714349129b55f322699c6770}'::uuid; + +-- Test if the UUID parser using SIMD optimization correctly rejects invalid UUID +-- string format. +SELECT '5b35380a714349129b55f32 99c6770'::uuid; +SELECT '5b35380a-7143-4912-9b55-f322699c67 '::uuid; +SELECT ' 35380a-7143-4912-9b55-f322699c6770'::uuid; +SELECT 'AZ35380a-7143-4912-9b55-f322699c6770'::uuid; +SELECT '{AZ35380a-7143-4912-9b55-f322699c6770}'::uuid; +SELECT '{AZ35380a714349129b55f322699c6770}'::uuid; +SELECT '{AZ35380a714349129b55f322699c67 }'::uuid; + +-- The parser only measures the input far enough to classify its shape. If it +-- stopped measuring at one of the accepted lengths, a longer string that +-- merely starts with a valid UUID would look like that UUID and be accepted +-- with the rest silently ignored, so check that trailing data is rejected. +SELECT '5b35380a714349129b55f322699c6770TRAILING'::uuid; +SELECT '{5b35380a714349129b55f322699c6770}TRAILING'::uuid; +SELECT '5b35380a-7143-4912-9b55-f322699c6770TRAILING'::uuid; +SELECT '{5b35380a-7143-4912-9b55-f322699c6770}TRAILING'::uuid; + -- clean up DROP TABLE guid1, guid2, guid3 CASCADE; -- 2.55.0
