I write: >From: ZizhuanLiu X-MAN <[email protected]> >Date: 2026-09-30 21:34 >To: pgsql-hackers <[email protected]> >Cc: Ilia Evdokimov <[email protected]>, tgl <[email protected]>, >tomas <[email protected]>, dean.a.rasheed <[email protected]>, >guofenglinux <[email protected]> >Subject: Re: Optimize MCV stats for sortable types and utilize sorted-order >properties > >Hi, > >1. Update the pg_stats_ext_exprs view to expose STATISTIC_KIND_MCV_VALUE_SORTED >through most_common_vals and most_common_freqs. > >2. Regarding pg_statistic_get_difference(): > >The input statistics can come either from pg_stats or directly from the user. >I initially considered improving the detection of the stat kind during import, >but there are several considerations and implementation difficulties. >For now, I think it is better to keep the existing logic. > >* Adding a stat kind parameter would rely on the caller to make the correct >decision, > and would also affect quite a few interfaces. >* Another idea is to inspect most_common_vals and most_common_freqs during >import, >sort them if possible, and generate STATISTIC_KIND_MCV_VALUE_SORTED; >otherwise, >keep STATISTIC_KIND_MCV. > >I looked into the latter, but it seems more complicated than expected because >both inputs are text. >We would need to deserialize them into the actual data type, sort them, and >serialize them again. >I have not worked out the details yet. If anyone is familiar with this part of >the code and has suggestions, >I would be happy to investigate further. > >3. I think pg_statistic_get_difference() could also have stronger validation >in the future. > >For input from pg_stats, the core code has already processed the statistics, >so for STATISTIC_KIND_MCV, >most_common_freqs[0] is the maximum and most_common_freqs[n - 1] is the >minimum. > >For user-provided input, however, there is currently no such restriction. We >only perform basic checks, >such as array lengths and NULLs. We could consider strengthening these checks >in the future.
Sorry, there was a mistake in my previous message. The function I meant was pg_restore_attribute_stats(), not pg_statistic_get_difference(). pg_restore_extended_stats() has the same issue: it always uses STATISTIC_KIND_MCV and does not check whether most_common_freqs[0] is the maximum frequency and most_common_freqs[n - 1] is the minimum frequency. One possible solution is to add an optional mcv_kind parameter to these two functions, with the default value being STATISTIC_KIND_MCV. This allows existing callers to keep the current behavior, while callers that provide STATISTIC_KIND_MCV_VALUE_SORTED statistics can explicitly preserve the correct statistics kind. This should minimize the impact on existing interfaces. Comments and suggestions are welcome. regards, -- ZizhuanLiu (X-MAN) [email protected]
