Patrick McFadin created CASSANDRA-21664:
-------------------------------------------

             Summary: Make schema changes cost O(1) in the number of tables and 
fail loudly at schema limits
                 Key: CASSANDRA-21664
                 URL: https://issues.apache.org/jira/browse/CASSANDRA-21664
             Project: Apache Cassandra
          Issue Type: Improvement
          Components: Cluster/Schema
            Reporter: Patrick McFadin


h2. What

Creating 10,000 tables takes 28 minutes (ManyTablesScalingTest, single node 
in-JVM). Per CREATE TABLE at small N: ~121 ms, of which 92 ms is a synchronous 
flush of every system_schema table on the log follower 
(SchemaKeyspace.applyChanges -> flush) and 22ms a blocking TCM log flush. On 
top of that, every schema change diffs the whole schema and rebuilds 
whole-keyspace maps, so cost grows with the number of existing tables: 112 -> 
234 ms per statement from 0 to 10,000 tables.

Past ~31,000 minimal tables the serialised cluster metadata exceeds 
max_mutation_size and the TCM snapshot silently stops being written (WARN 
only). Table-count guardrails ship disabled.

h2. Change

* Schema diffs and keyspace/table map updates touch only what changed: 
Keyspaces.diff, Tables (persistent map), DistributedSchema table map, 
Keyspaces.withAddedOrUpdated, DistributedSchema.validate, ThreadLocalMeter 
array growth.
* system_schema flush is coalesced behind a new yaml setting 
schema_flush_coalescing_window (default 1000ms; 0ms restores the synchronous 
flush). system_schema has durable_writes and is rebuilt from the cluster 
metadata log on startup, so delaying the flush cannot lose schema. Drain still 
flushes synchronously.
* Snapshot store failure is logged at ERROR with a metric 
(TCM.SnapshotStoreFailures), and a WARN fires once serialised metadata passes 
75% of max_mutation_size.
* tables_warn_threshold defaults to 1000; tested envelope and per-table cost 
documented.
* Build: the in-JVM dtest targets' -Xmx8G was overridden by junit's maxmemory 
default of 1024m.

h2. Result

Same test, same machine; before = synchronous flush, after = this patch.
||N||before||after||per statement, first -> last bucket (after)||
|1,000|120.9 s|29.9 s|32 -> 31 ms|
|5,000|675.1 s|204.5 s|32 -> 55 ms|
|10,000|28.0 min|9.7 min|32 -> 95 ms|

The fixed per-statement floor is gone. A residual ~6 ms per 1,000 existing 
tables remains (TablesDiff full scans); follow-up.

h2. Tests

Scaling assertions at N vs 8N on every touched path; SchemaFlushCoalesceTest; 
SchemaFlushRestartTest (30 tables, stop without drain, restart, all present); 
MetadataSnapshotSizeWarningTest; schema, tcm, guardrail and config suites; 
checkstyle clean.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to