[
https://issues.apache.org/jira/browse/CASSANDRA-21664?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Patrick McFadin reassigned CASSANDRA-21664:
-------------------------------------------
Assignee: Patrick McFadin
> Make schema changes cost O(1) in the number of tables and fail loudly at
> schema limits
> --------------------------------------------------------------------------------------
>
> Key: CASSANDRA-21664
> URL: https://issues.apache.org/jira/browse/CASSANDRA-21664
> Project: Apache Cassandra
> Issue Type: Improvement
> Components: Cluster/Schema
> Reporter: Patrick McFadin
> Assignee: Patrick McFadin
> Priority: Normal
>
> h2. What
> Creating 10,000 tables takes 28 minutes (ManyTablesScalingTest, single node
> in-JVM). Per CREATE TABLE at small N: ~121 ms, of which 92 ms is a
> synchronous flush of every system_schema table on the log follower
> (SchemaKeyspace.applyChanges -> flush) and 22ms a blocking TCM log flush. On
> top of that, every schema change diffs the whole schema and rebuilds
> whole-keyspace maps, so cost grows with the number of existing tables: 112 ->
> 234 ms per statement from 0 to 10,000 tables.
> Past ~31,000 minimal tables the serialised cluster metadata exceeds
> max_mutation_size and the TCM snapshot silently stops being written (WARN
> only). Table-count guardrails ship disabled.
> h2. Change
> * Schema diffs and keyspace/table map updates touch only what changed:
> Keyspaces.diff, Tables (persistent map), DistributedSchema table map,
> Keyspaces.withAddedOrUpdated, DistributedSchema.validate, ThreadLocalMeter
> array growth.
> * system_schema flush is coalesced behind a new yaml setting
> schema_flush_coalescing_window (default 1000ms; 0ms restores the synchronous
> flush). system_schema has durable_writes and is rebuilt from the cluster
> metadata log on startup, so delaying the flush cannot lose schema. Drain
> still flushes synchronously.
> * Snapshot store failure is logged at ERROR with a metric
> (TCM.SnapshotStoreFailures), and a WARN fires once serialised metadata passes
> 75% of max_mutation_size.
> * tables_warn_threshold defaults to 1000; tested envelope and per-table cost
> documented.
> * Build: the in-JVM dtest targets' -Xmx8G was overridden by junit's maxmemory
> default of 1024m.
> h2. Result
> Same test, same machine; before = synchronous flush, after = this patch.
> ||N||before||after||per statement, first -> last bucket (after)||
> |1,000|120.9 s|29.9 s|32 -> 31 ms|
> |5,000|675.1 s|204.5 s|32 -> 55 ms|
> |10,000|28.0 min|9.7 min|32 -> 95 ms|
> The fixed per-statement floor is gone. A residual ~6 ms per 1,000 existing
> tables remains (TablesDiff full scans); follow-up.
> h2. Tests
> Scaling assertions at N vs 8N on every touched path; SchemaFlushCoalesceTest;
> SchemaFlushRestartTest (30 tables, stop without drain, restart, all present);
> MetadataSnapshotSizeWarningTest; schema, tcm, guardrail and config suites;
> checkstyle clean.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]