Hi all,

Following up on HIVE-30052/30053 (CachedStore prewarm), we found another
metastore scalability issue while preparing our HMS 2.3.6 to 4.2.0
migration: HIVE-30096.

For any non-rename, table-level alter of an unpartitioned table,
HiveAlterHandler calls MetaStoreServerUtils.updateTableStatsSlow with
forceRecompute=true unconditionally. The metastore then recursively lists
the entire table location and materializes one FileStatus per file in
memory, inside a single RPC. No server-side config gates this path
(metastore.stats.autogather is only consulted on the create/add-partition
paths), and the same code exists at least back to 2.3.x.

This is severe for table formats that keep their partitioning in a
transaction log (e.g. Delta, Iceberg): they register as unpartitioned
tables with millions of files under one directory. On our deployment, a
property-only alter of a Delta table with ~4M files on GCS materialized
~7.9GB of FileStatus objects and OOMed the metastore with an 8GB heap.
Client behavior varies: Trino always sends DO_NOT_UPDATE_STATS=true on
alters (its source comments describe exactly this problem), while Spark's
built-in Hive client (through at least 3.5) sends no EnvironmentContext at
all (an old TODO in its HiveShim references HIVE-12730), so every Spark
alter takes this path.

The fix in https://github.com/apache/hive/pull/6822 forces the recompute
only when the alter actually changed the table location. Metadata-only
alters then reuse the fast stats already present via the existing
containsAllFastStats shortcut; location changes recompute as before; tables
missing the fast stats are still listed once and self-heal. This also makes
the alter path consistent with the create path, which has always passed
forceRecompute=false.

A broader question for the list, mostly about consumers: for unpartitioned
tables, is anyone aware of readers that depend specifically on the
alter-time refresh of the fast stats (numFiles/totalSize), as opposed to
the values written at create time, by StatsTask on writes, or by ANALYZE?
The alter-time recompute marks the result as not-accurate
(setBasicStatsState(FALSE)) for any non-ANALYZE alter, so as far as we can
tell the CBO paths never trust the numbers this particular listing produces
but if there are consumers we have not considered, that would be useful
input on #6822.

I would also like input on a follow-up idea: making the quick-stats
computation itself stream file statuses (RemoteIterator, O(1) memory)
instead of materializing a List<FileStatus>, which would bound metastore
memory on the remaining legitimate listing paths (create over an existing
directory, location changes). Happy to file a JIRA and implement if there
is interest.

Reviews of #6822 would be much appreciated.

Thanks,
Vidit Gupta
Data Platform, Meesho

-- 
***
This communication is confidential, may be privileged, and is meant 
only for the intended recipient and purpose. No part of this email or any 
files transmitted with it can be shared, copied, forwarded, published 
online or offline, or used in any unauthorised manner. If you are not the 
intended recipient, please preserve the confidentiality of the contents, 
delete the e-mail and attachments (if any) from your system, and inform the 
sender immediately.
***

Reply via email to