[ 
https://issues.apache.org/jira/browse/CASSANDRA-21664?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Patrick McFadin reassigned CASSANDRA-21664:
-------------------------------------------

    Assignee: Patrick McFadin

> Make schema changes cost O(1) in the number of tables and fail loudly at 
> schema limits
> --------------------------------------------------------------------------------------
>
>                 Key: CASSANDRA-21664
>                 URL: https://issues.apache.org/jira/browse/CASSANDRA-21664
>             Project: Apache Cassandra
>          Issue Type: Improvement
>          Components: Cluster/Schema
>            Reporter: Patrick McFadin
>            Assignee: Patrick McFadin
>            Priority: Normal
>
> h2. What
> Creating 10,000 tables takes 28 minutes (ManyTablesScalingTest, single node 
> in-JVM). Per CREATE TABLE at small N: ~121 ms, of which 92 ms is a 
> synchronous flush of every system_schema table on the log follower 
> (SchemaKeyspace.applyChanges -> flush) and 22ms a blocking TCM log flush. On 
> top of that, every schema change diffs the whole schema and rebuilds 
> whole-keyspace maps, so cost grows with the number of existing tables: 112 -> 
> 234 ms per statement from 0 to 10,000 tables.
> Past ~31,000 minimal tables the serialised cluster metadata exceeds 
> max_mutation_size and the TCM snapshot silently stops being written (WARN 
> only). Table-count guardrails ship disabled.
> h2. Change
> * Schema diffs and keyspace/table map updates touch only what changed: 
> Keyspaces.diff, Tables (persistent map), DistributedSchema table map, 
> Keyspaces.withAddedOrUpdated, DistributedSchema.validate, ThreadLocalMeter 
> array growth.
> * system_schema flush is coalesced behind a new yaml setting 
> schema_flush_coalescing_window (default 1000ms; 0ms restores the synchronous 
> flush). system_schema has durable_writes and is rebuilt from the cluster 
> metadata log on startup, so delaying the flush cannot lose schema. Drain 
> still flushes synchronously.
> * Snapshot store failure is logged at ERROR with a metric 
> (TCM.SnapshotStoreFailures), and a WARN fires once serialised metadata passes 
> 75% of max_mutation_size.
> * tables_warn_threshold defaults to 1000; tested envelope and per-table cost 
> documented.
> * Build: the in-JVM dtest targets' -Xmx8G was overridden by junit's maxmemory 
> default of 1024m.
> h2. Result
> Same test, same machine; before = synchronous flush, after = this patch.
> ||N||before||after||per statement, first -> last bucket (after)||
> |1,000|120.9 s|29.9 s|32 -> 31 ms|
> |5,000|675.1 s|204.5 s|32 -> 55 ms|
> |10,000|28.0 min|9.7 min|32 -> 95 ms|
> The fixed per-statement floor is gone. A residual ~6 ms per 1,000 existing 
> tables remains (TablesDiff full scans); follow-up.
> h2. Tests
> Scaling assertions at N vs 8N on every touched path; SchemaFlushCoalesceTest; 
> SchemaFlushRestartTest (30 tables, stop without drain, restart, all present); 
> MetadataSnapshotSizeWarningTest; schema, tcm, guardrail and config suites; 
> checkstyle clean.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to