Fix potential buffer overrun in regexp match/split functions.

setup_regexp_matches() sizes the buffer used to convert matched
substrings back from pg_wchar form at the smaller of maxlen*eml and
the original string's byte length, on the assumption that such a
conversion cannot produce more bytes than the string it came
from. That assumption holds only for validly encoded input. But
pg_mb2wchar_with_len() silently accepts bytes that are invalid in the
database encoding, turning each such byte into one pg_wchar, and
converting that back can take more bytes than the input did. A string
made of such bytes therefore overruns the conversion buffer by up to
its own length, corrupting the following memory. regexp_match(),
regexp_matches(), regexp_split_to_table() and regexp_split_to_array()
are all affected.

Fix by dropping the tighter bound and always allocating maxlen*eml + 1
bytes.

Reported-by: Francesco Verardi <[email protected]>
Author: Masahiko Sawada <[email protected]>
Reviewed-by: Tom Lane <[email protected]>
Backpatch-through: 14
Security: CVE-2026-14664

Branch
------
master

Details
-------
https://git.postgresql.org/pg/commitdiff/343aabf1d95dfb65c0b7c7cdc2e2fb4e79ba9e6d
Author: Masahiko Sawada <[email protected]>

Modified Files
--------------
src/backend/utils/adt/regexp.c | 23 ++++++++++++-----------
1 file changed, 12 insertions(+), 11 deletions(-)

Reply via email to