Nice patch, I like the test. a few comments below On Wed, Sep 30, 2026 at 11:43 PM Rosemarie O'Riorden via dev < [email protected]> wrote:
> This commit adds a test that developers can run to check if the changes > they've made cause a regression in the performance of ovn-controller or > ovn-northd. > > Run from inside the sandbox and compare two commits' runtime and peak > memory usage for each process. > > This test creates a topology consisting of logical switches and routers, > ACLs, port groups, NAT, DHCP, DNS, QoS, routing policies, and load > balancers. The user can also provide their own database file for the > benchmark. > > * To run the test with defaults (200 nodes, track northd and controller): > ./ovn-benchmark.sh > * To run with 50 nodes: ./ovn-benchmark.sh 50 > * To run with 30 nodes and only track ovn-northd: > ./ovn-benchmark.sh 30 northd > * To run with Valgrind Massif tracking allocations: > ./ovn-benchmark.sh --valgrind > * To run with a custom db file: ./ovn-benchmark.sh -f file.db > * To run with debug info printed: ./ovn-benchmark.sh --debug > * To see detailed usage instructions: ./ovn-benchmark.sh --help > > Each "node" is a router-switch pair with associated features: > - 1 gateway router (with NAT, static routes, routing policies) > - 1 logical switch (with 9 ports by default, DHCP, DNS, QoS) > - 5 load balancers (each with 5 backends) > - ACLs, port groups, and address sets (shared across all nodes) > > For example, 200 nodes creates: > - 200 routers + 200 switches > - 1,800 logical switch ports (9 per switch) > - 1,000 load balancers (5 per node) > - 600 NAT rules, 600 static routes, 400 routing policies > nit: Was this updated when you added IPv6 support? for NAT rules it looks like each node gets an SNAT, DNAT, and DNAT_AND_SNAT for IPv4 and IPv6. This would mean 200 nodes generate 1200 NAT rules. I think it is a similar change for static routes (except each node gets a route_specific, route_insert, and route_discard for IPv4 and IPv6). > - Plus ACLs, DHCP/DNS records, QoS rules, etc. > > ovn-benchmark.py creates the topology and ovn-benchmark.sh is the > wrapper that tracks peak memory and execution time. ovn-benchmark.py was > built off of ovn-lb-benchmark.py as a starting point. > > VmPeak is tracked by default, but Valgrind can be used instead for more > precise results. > > Reported-at: https://redhat.atlassian.net/browse/FDP-978 > Assisted-by: Claude Sonnet 4.5, Claude Code > Signed-off-by: Rosemarie O'Riorden <[email protected]> > --- > BTW: > A CI component for this patch, which would add a GH action that runs the > benchmark on the base upstream vs with the patch applied, and compares > the results to test for regressions, is coming in the future. > > v2 -> v3: > - Memory tracking: > - Dropped the background watcher loop that read RSS periodically. > Now VmPeak is read once from /proc/<pid>/status instead, so short > peaks aren't missed. > - Added --valgrind (-v) to rerun the tracked processes under Massif > and report peak heap. More accurate but 10-50x slower, so it > warns above 500 nodes. > - The processes valgrind hijacks get restarted on exit, Ctrl-C > included, from their saved cmdline and cwd. > - Timing: > - Total time is split into topology generation, OVN sync and db > compaction, so you can tell which phase regressed. > - Topology fixes: > - DHCP options were created but never attached to any port. They're > assigned now, v4 and v6. > - Same for the web/db/app address sets -- nothing referenced them, so > added ACLs that do. > - Split trusted_networks into v4 and v6 sets; one mixed set can't > serve both the ip4.dst and ip6.dst matches. > - Gateway join ports all shared one MAC and 10.0.0.1/8. Each gets > its own MAC and a real join subnet address now, + routes to match. > - The port group drop ACL matches outport instead of inport. > - Chassis names start at 1 to match the sandbox system-id. > - Safety: > - A missing router, switch, or port group used to be skipped quietly, > leaving a half-built topology. These error out now. > - Errors out on duplicate processe and script running outside the > sandbox. > - Added a SIGINT trap, bounds checks for --backends and --vips, and a > fix for binding more ports than there are nodes. > - Docs: > - testing.rst covers both memory modes and the valgrind tradeoff. > - Small stuff: > - Removed an unused json RPC connection. > - Port group lookups use a dict instead of rescanning the table. > - Better warnings for misused args, and options that may be confusing. > - Quoting, output formatting, comments, etc. > > v1 -> v2: > - Topology: > - Restructured from 1:1 chassis:node to a nodes-per-chassis style > where each chassis handles a partition or 'batch' of the network > (ascii > diagram added). > - Added --batch-size (-b) for nodes-per-chassis grouping. > - Added IPv6 coverage. > - IP addressing: > - ip_node()/ip6_node() helpers spread index across two octets to support > >255 nodes. > - Added bounds checks for node count and ports-per-switch. > - Shell fixes: > - pgrep -x for exact process match. > - break 2 to exit both watcher loops on process death. > - trap cleanup EXIT with mktemp -d for temp files. > - Renamed SECONDS bc it's a bash keyword. > - python -> python3. > - Quoted array expansions, fixed PID error messages. > - Code cleanup: > - Extracted _add_node(), _add_ports(), _add_acl() helpers. > - Removed dead code (find_by_name(), seqno). > - Fixed error variable shadowing imported module 'error'. > - Docs: > - Added ovn-benchmark section to testing.rst. > - Added note in submitting-patches.rst. > --- > .../contributing/submitting-patches.rst | 4 + > Documentation/topics/testing.rst | 37 + > tutorial/automake.mk | 4 +- > tutorial/ovn-benchmark.py | 905 ++++++++++++++++++ > tutorial/ovn-benchmark.sh | 467 +++++++++ > 5 files changed, 1416 insertions(+), 1 deletion(-) > create mode 100755 tutorial/ovn-benchmark.py > create mode 100755 tutorial/ovn-benchmark.sh > > diff --git a/Documentation/internals/contributing/submitting-patches.rst > b/Documentation/internals/contributing/submitting-patches.rst > index abeb3e9d0..c96fdff84 100644 > --- a/Documentation/internals/contributing/submitting-patches.rst > +++ b/Documentation/internals/contributing/submitting-patches.rst > @@ -68,6 +68,10 @@ Testing is also important: > feature. A bug fix patch should preferably add a test that would > fail if the bug recurs. > > +- To check for memory or performance regressions, you can run > + ``./ovn-benchmark.sh`` inside the OVN sandbox (``make sandbox``). > + See the "ovn-benchmark" section of :doc:`/topics/testing` for details. > + > If you are using GitHub, then you may utilize the GitHub Actions CI > system. > This will run the above tests automatically when you push changes to your > repository. > diff --git a/Documentation/topics/testing.rst > b/Documentation/topics/testing.rst > index 951133c29..da5046d2b 100644 > --- a/Documentation/topics/testing.rst > +++ b/Documentation/topics/testing.rst > @@ -294,6 +294,43 @@ of these cached objects, be sure to rebuild the test. > The cached objects are stored under the relevant folder in > ``tests/perf-testsuite.dir/cached``. > > +ovn-benchmark > ++++++++++++++ > + > +The ``tutorial/`` directory contains a memory and performance benchmarking > +tool that can be used to detect regressions between commits: > + > +- ``ovn-benchmark.sh``: Shell wrapper that tracks peak memory and > + execution time for ``ovn-northd`` and ``ovn-controller``. By default > + it reports peak virtual memory (VmPeak) from ``/proc/<pid>/status``, > + which captures all allocated memory including pages not yet accessed. > + Pass ``--valgrind`` to restart the tracked processes under Valgrind > + Massif and report peak heap memory instead; this is more accurate but > + 10-50x slower. > +- ``ovn-benchmark.py``: Python script that populates the Northbound > database > + with a realistic ovn-kubernetes-style topology. > + > +The generated topology includes gateway routers, logical switches, NAT > rules, > +ACLs, DHCP/DHCPv6, DNS, load balancers, QoS rules, address sets, port > groups, > +static routes, and routing policies -- all with both IPv4 and IPv6 > coverage. > + > +The benchmark must be run from inside the OVN sandbox. Run > +``./ovn-benchmark.sh --help`` for the full list of options:: > + > + $ make sandbox > + $ ./ovn-benchmark.sh > + > +By default, the topology is generated through many individual OVSDB > +transactions via ``ovn-benchmark.py``. > + > +.. note:: > + > + The ``-f`` option loads a Northbound database file via ``ovsdb-client > + restore`` instead, which applies the entire database as a single > + transaction. This may yield different memory and timing results > because > + ``ovn-northd`` and ``ovn-controller`` process one large batch of > changes > + rather than reacting to each transaction incrementally. > + > OVN Upgrade Testing > ~~~~~~~~~~~~~~~~~~~ > > diff --git a/tutorial/automake.mk b/tutorial/automake.mk > index 631208639..fdea2449c 100644 > --- a/tutorial/automake.mk > +++ b/tutorial/automake.mk > @@ -2,7 +2,9 @@ EXTRA_DIST += \ > tutorial/ovn-sandbox \ > tutorial/ovn-setup.sh \ > tutorial/ovn-lb-benchmark.sh \ > - tutorial/ovn-lb-benchmark.py > + tutorial/ovn-lb-benchmark.py \ > + tutorial/ovn-benchmark.sh \ > + tutorial/ovn-benchmark.py > sandbox: all > cd $(srcdir)/tutorial && MAKE=$(MAKE) HAVE_OPENSSL=$(HAVE_OPENSSL) > \ > ./ovn-sandbox -b $(abs_builddir) --ovs-src $(ovs_srcdir) > --ovs-build $(ovs_builddir) $(SANDBOXFLAGS) > diff --git a/tutorial/ovn-benchmark.py b/tutorial/ovn-benchmark.py > new file mode 100755 > index 000000000..0d48778c8 > --- /dev/null > +++ b/tutorial/ovn-benchmark.py > @@ -0,0 +1,905 @@ > +#!/usr/bin/env python3 > +"""OVN memory regression testing tool. > + > +Creates a broad OVN topology to detect memory regressions between commits. > +Designed to be run via ovn-benchmark.sh. (Run ./ovn-benchmark.sh --help > +to see usage). > + > +Topology (simulates the default ovn-kubernetes topology): > + > + lsp-0-* lsp-1-* (Workload ports) > + | | > + ls-0 ls-1 (Logical Switches) > + | | > + s2c-0/c2s-0 s2c-1/c2s-1 > + \\ / > + +--- cluster (LR) ---+ (Cluster Router) > + | > + sjc/rcj > + | > + +---- join (LS) ----+ (Join Switch) > + / \\ > + j2lr-0/lr2j-0 j2lr-1/lr2j-1 > + | | > + lr-0 lr-1 (Gateway Routers) > + > + Chassis binding (with batch size B, default n/10): > + chassis-1: c2s-[0..B), lr-[0..B) (matches sandbox system-id) > + chassis-2: c2s-[B..2B), lr-[B..2B) > + ... > + > +Each node creates a dual-stack gateway router + logical switch pair > +with: > + - Dual-stack addressing (IPv4 10.x.x.x/24, IPv6 fd00:x::/64) > + - NAT (SNAT/DNAT/DNAT_AND_SNAT), static routes, routing policies > + - Configurable ports per switch with port security > + - Security: Address sets, port groups, ACLs > + - Services: DHCP, DHCPv6, DNS, load balancers > + - QoS: Bandwidth limiting, DSCP marking > + > +Note: Uses explicit (non-templated) load balancers to maximize memory > +usage for regression testing. For templated LB testing, see > +ovn-lb-benchmark.py. > +""" > + > +import argparse > +import sys > + > +import ovs.db.idl > +import ovs.poller > +import ovs.vlog > +from ovs.db import error > + > +vlog = ovs.vlog.Vlog('ovn-benchmark') > +vlog.set_levels_from_string('console:warn') > +vlog.init(None) > + > +SCHEMA = '../ovn-nb.ovsschema' > + > + > +def die(msg): > + sys.stderr.write(f'\nError: {msg}\n') > + sys.exit(1) > + > + > +def ip_node(i): > + """Convert node index to two IP octets, supporting up to 65535 > nodes.""" > + return f'{i >> 8}.{i & 0xff}' > + > + > +def ip6_node(i): > + """Convert node index to an IPv6 hextet for ULA addresses.""" > + return f'{i:x}' > + > + > +def create_address_sets(idl, n): > + """Create address sets for security groups.""" > + vlog.info('Creating address sets') > + txn = ovs.db.idl.Transaction(idl) > + > + web_as = txn.insert(idl.tables['Address_Set']) > + web_as.name = 'web_servers' > + web_as.addresses = ( > + [f'10.{ip_node(i)}.10' for i in range(n)] > + + [f'fd00:{ip6_node(i)}::a' for i in range(n)]) > + > + db_as = txn.insert(idl.tables['Address_Set']) > + db_as.name = 'db_servers' > + db_as.addresses = ( > + [f'10.{ip_node(i)}.20' for i in range(n)] > + + [f'fd00:{ip6_node(i)}::14' for i in range(n)]) > + > + app_as = txn.insert(idl.tables['Address_Set']) > + app_as.name = 'app_servers' > + app_as.addresses = ( > + [f'10.{ip_node(i)}.30' for i in range(n)] > + + [f'fd00:{ip6_node(i)}::1e' for i in range(n)]) > + > + trusted_as_v4 = txn.insert(idl.tables['Address_Set']) > + trusted_as_v4.name = 'trusted_networks_v4' > + trusted_as_v4.addresses = ['192.168.0.0/16', '172.16.0.0/12'] > + > + trusted_as_v6 = txn.insert(idl.tables['Address_Set']) > + trusted_as_v6.name = 'trusted_networks_v6' > + trusted_as_v6.addresses = ['fc00::/7'] > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to create address sets ({txn.get_error()})') > + > + > +def create_port_groups(idl): > + """Create port groups for security group implementation.""" > + vlog.info('Creating port groups') > + txn = ovs.db.idl.Transaction(idl) > + > + web_pg = txn.insert(idl.tables['Port_Group']) > + web_pg.name = 'web_tier' > + > + db_pg = txn.insert(idl.tables['Port_Group']) > + db_pg.name = 'db_tier' > + > + app_pg = txn.insert(idl.tables['Port_Group']) > + app_pg.name = 'app_tier' > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to create port groups ({txn.get_error()})') > + > + > +def create_dhcp_options(idl, n): > + """Create DHCP options for each subnet.""" > + for i in range(n): > + vlog.info(f'Creating DHCP options for node {i}') > + txn = ovs.db.idl.Transaction(idl) > + dhcp_opts = txn.insert(idl.tables['DHCP_Options']) > + dhcp_opts.cidr = f'10.{ip_node(i)}.0/24' > + dhcp_opts.setkey('options', 'server_id', f'10.{ip_node(i)}.1') > + dhcp_opts.setkey('options', 'server_mac', '00:00:00:00:00:01') > + dhcp_opts.setkey('options', 'lease_time', '3600') > + dhcp_opts.setkey('options', 'router', f'10.{ip_node(i)}.1') > + dhcp_opts.setkey('options', 'dns_server', f'10.{ip_node(i)}.2') > + dhcp_opts.setkey('options', 'domain_name', '"example.com"') > + dhcp_opts.setkey('options', 'mtu', '1500') > + dhcp_opts.setkey('external_ids', 'subnet', f'ls-{i}') > + > + dhcpv6_opts = txn.insert(idl.tables['DHCP_Options']) > + dhcpv6_opts.cidr = f'fd00:{ip6_node(i)}::/64' > + dhcpv6_opts.setkey('options', 'server_id', > + '00:00:00:00:00:01') > + dhcpv6_opts.setkey('options', 'dns_server', > + f'fd00:{ip6_node(i)}::2') > + dhcpv6_opts.setkey('options', 'domain_search', > + '"example.com"') > + dhcpv6_opts.setkey('external_ids', 'subnet', > + f'ls-{i}-v6') > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to create DHCP options for node {i} ' > + f'({txn.get_error()})') > + > + > +def assign_dhcp_options(idl, n): > + """Assign DHCP options to logical switch ports.""" > + dhcp4 = {} > + dhcp6 = {} > + for row in idl.tables['DHCP_Options'].rows.values(): > + subnet = row.external_ids.get('subnet', '') > + if subnet.endswith('-v6'): > + dhcp6[subnet.removesuffix('-v6')] = row str.removesuffix() is introduced in python 3.9 according to OVN_CHECK_PYTHON3 the earliest version of python that can be used is 3.7. can this be changed to subnet[:-3]? > + else: > + dhcp4[subnet] = row > + > + for i in range(n): > + vlog.info(f'Assigning DHCP options for node {i}') > + txn = ovs.db.idl.Transaction(idl) > + opts4 = dhcp4.get(f'ls-{i}') > + opts6 = dhcp6.get(f'ls-{i}') > + for row in idl.tables['Logical_Switch_Port'].rows.values(): > + if row.name.startswith(f'lsp-{i}-'): > + if opts4: > + row.dhcpv4_options = opts4.uuid > + if opts6: > + row.dhcpv6_options = opts6.uuid > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to assign DHCP options for node {i} ' > + f'({txn.get_error()})') > + > + > +def create_qos_rules(idl, n, switches): > + """Create QoS rules for bandwidth limiting and DSCP marking.""" > + for i in range(n): > + vlog.info(f'Creating QoS rules for node {i}') > + txn = ovs.db.idl.Transaction(idl) > + > + ls = switches.get(f'ls-{i}') > + if not ls: > + die(f'Missing switch ls-{i}') > + > + qos_bw = txn.insert(idl.tables['QoS']) > + qos_bw.priority = 100 > + qos_bw.direction = 'to-lport' > + qos_bw.match = f'inport == "lsp-{i}-0"' > + qos_bw.setkey('bandwidth', 'rate', 1000) > + qos_bw.setkey('bandwidth', 'burst', 100) > + qos_bw.setkey('external_ids', 'type', 'rate-limit') > + ls.addvalue('qos_rules', qos_bw.uuid) > + > + qos_dscp = txn.insert(idl.tables['QoS']) > + qos_dscp.priority = 200 > + qos_dscp.direction = 'from-lport' > + qos_dscp.match = 'ip4 && tcp.dst == 22' > + qos_dscp.setkey('action', 'dscp', 46) > + qos_dscp.setkey('external_ids', 'type', 'dscp-marking') > + ls.addvalue('qos_rules', qos_dscp.uuid) > + > + qos_mark = txn.insert(idl.tables['QoS']) > + qos_mark.priority = 150 > + qos_mark.direction = 'from-lport' > + qos_mark.match = 'ip4 && udp' > + qos_mark.setkey('action', 'mark', 1) > + qos_mark.setkey('external_ids', 'type', 'packet-marking') > + ls.addvalue('qos_rules', qos_mark.uuid) > + > + qos_dscp_v6 = txn.insert(idl.tables['QoS']) > + qos_dscp_v6.priority = 200 > + qos_dscp_v6.direction = 'from-lport' > + qos_dscp_v6.match = 'ip6 && tcp.dst == 22' > + qos_dscp_v6.setkey('action', 'dscp', 46) > + qos_dscp_v6.setkey('external_ids', 'type', > + 'dscp-marking-v6') > + ls.addvalue('qos_rules', qos_dscp_v6.uuid) > + > + qos_mark_v6 = txn.insert(idl.tables['QoS']) > + qos_mark_v6.priority = 150 > + qos_mark_v6.direction = 'from-lport' > + qos_mark_v6.match = 'ip6 && udp' > + qos_mark_v6.setkey('action', 'mark', 1) > + qos_mark_v6.setkey('external_ids', 'type', > + 'packet-marking-v6') > + ls.addvalue('qos_rules', qos_mark_v6.uuid) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to create QoS rules for node {i} > ({txn.get_error()})') > + > + > +def create_acls_for_port_group(idl, port_groups, pg_name, allowed_ports): > + """Create ACLs for a specific port group.""" > + txn = ovs.db.idl.Transaction(idl) > + > + pg = port_groups.get(pg_name) > + if not pg: > + die(f'Missing port group {pg_name}') > + > + acl_allow_est = txn.insert(idl.tables['ACL']) > + acl_allow_est.priority = 1100 > + acl_allow_est.direction = 'to-lport' > + acl_allow_est.match = 'ct.est && !ct.rel && !ct.new && !ct.inv' > + acl_allow_est.action = 'allow-related' > + pg.addvalue('acls', acl_allow_est.uuid) > + > + acl_allow_rel = txn.insert(idl.tables['ACL']) > + acl_allow_rel.priority = 1100 > + acl_allow_rel.direction = 'to-lport' > + acl_allow_rel.match = 'ct.rel && !ct.est && !ct.new && !ct.inv' > + acl_allow_rel.action = 'allow-related' > + pg.addvalue('acls', acl_allow_rel.uuid) > + > + for port in allowed_ports: > + acl_new = txn.insert(idl.tables['ACL']) > + acl_new.priority = 1050 > + acl_new.direction = 'to-lport' > + acl_new.match = f'ct.new && tcp.dst == {port}' > + acl_new.action = 'allow-related' > + pg.addvalue('acls', acl_new.uuid) > + > + acl_drop = txn.insert(idl.tables['ACL']) > + acl_drop.priority = 1000 > + acl_drop.direction = 'to-lport' > + acl_drop.match = f'outport == @{pg_name}' > + acl_drop.action = 'drop' > + pg.addvalue('acls', acl_drop.uuid) > + > + acl_arp = txn.insert(idl.tables['ACL']) > + acl_arp.priority = 1010 > + acl_arp.direction = 'to-lport' > + acl_arp.match = 'arp || nd' > + acl_arp.action = 'allow' > + pg.addvalue('acls', acl_arp.uuid) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to create ACLs for {pg_name} ({txn.get_error()})') > + > + > +def create_acls(idl): > + """Create comprehensive ACLs for port groups. > + > + ACL Priority Allocation: > + 2500-2999: Security enforcement (anti-spoofing, etc.) > + 2000-2499: Management access (SSH, ICMP, DHCP) > + 1000-1499: Port group ACLs (security groups) > + 1100: Connection tracking (established, related) > + 1050: New connection per-port rules > + 1010: ARP/ND allow > + 1000: Default drop > + > + Switch ACLs (higher priority) can override port-group ACLs, with > + security enforcement (anti-spoofing) taking highest priority. > + """ > + vlog.info('Creating ACLs for port groups') > + port_groups = {row.name: row > + for row in idl.tables['Port_Group'].rows.values()} > + create_acls_for_port_group(idl, port_groups, 'web_tier', [80, 443]) > + create_acls_for_port_group(idl, port_groups, 'app_tier', [8080, 9000]) > + create_acls_for_port_group(idl, port_groups, 'db_tier', [5432, 3306]) > + > + > +def _add_acl(txn, idl, ls, priority, direction, match, action, > + description=None): > + """Create an ACL and attach it to a logical switch.""" > + acl = txn.insert(idl.tables['ACL']) > + acl.priority = priority > + acl.direction = direction > + acl.match = match > + acl.action = action > + if description: > + acl.setkey('external_ids', 'description', description) > + ls.addvalue('acls', acl.uuid) > + > + > +def add_acls_to_switch(idl, switch_name, node_id, switches): > + """Add ACLs directly to a logical switch. > + > + Adds switch-level ACLs for SSH, ICMP, anti-spoofing, and DHCP. > + Anti-spoofing (priority 2500) prevents VMs from using IPs outside > + their assigned subnet, blocking IP address spoofing attacks. > + """ > + txn = ovs.db.idl.Transaction(idl) > + > + ls = switches.get(switch_name) > + if not ls: > + die(f'Missing switch {switch_name}') > + > + _add_acl(txn, idl, ls, 2000, 'from-lport', > + 'tcp.dst == 22 && ip4.src == 10.0.0.0/8', > + 'allow', 'Allow SSH from internal') > + _add_acl(txn, idl, ls, 1500, 'from-lport', > + 'icmp4 || icmp6', 'allow') > + _add_acl(txn, idl, ls, 2500, 'from-lport', > + f'ip4.src != 10.{ip_node(node_id)}.0/24', > + 'drop', 'Anti-spoofing') > + _add_acl(txn, idl, ls, 2000, 'from-lport', > + 'udp.src == 68 && udp.dst == 67', 'allow') > + > + _add_acl(txn, idl, ls, 2000, 'from-lport', > + 'tcp.dst == 22 && ip6.src == fd00::/16', > + 'allow', 'Allow SSH from internal (IPv6)') > + _add_acl(txn, idl, ls, 2500, 'from-lport', > + f'ip6.src != fd00:{ip6_node(node_id)}::/64', > + 'drop', 'Anti-spoofing (IPv6)') > + _add_acl(txn, idl, ls, 2000, 'from-lport', > + 'udp.src == 546 && udp.dst == 547', 'allow') > + > + _add_acl(txn, idl, ls, 1800, 'to-lport', > + 'ip4.dst == $web_servers && tcp.dst == 80', > + 'allow-related', 'Allow HTTP to web servers') > + _add_acl(txn, idl, ls, 1800, 'to-lport', > + 'ip4.dst == $db_servers && tcp.dst == 5432', > + 'allow-related', 'Allow Postgres to DB servers') > + _add_acl(txn, idl, ls, 1800, 'to-lport', > + 'ip4.dst == $app_servers && tcp.dst == 8080', > + 'allow-related', 'Allow HTTP to app servers') > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to add ACLs to switch {switch_name} ' > + f'({txn.get_error()})') > + > + > +def create_dns_records(idl, n, switches): > + """Create DNS records in the NB database.""" > + for i in range(n): > + vlog.info(f'Creating DNS records for node {i}') > + txn = ovs.db.idl.Transaction(idl) > + dns = txn.insert(idl.tables['DNS']) > + dns.setkey('records', f'web-{i}.example.com', > + f'10.{ip_node(i)}.10 ' > + f'fd00:{ip6_node(i)}::a') > + dns.setkey('records', f'app-{i}.example.com', > + f'10.{ip_node(i)}.30 ' > + f'fd00:{ip6_node(i)}::1e') > + dns.setkey('records', f'db-{i}.example.com', > + f'10.{ip_node(i)}.20 ' > + f'fd00:{ip6_node(i)}::14') > + dns.setkey('external_ids', 'zone', f'zone-{i}') > + > + ls = switches.get(f'ls-{i}') > + if not ls: > + die(f'Missing switch ls-{i}') > + ls.addvalue('dns_records', dns.uuid) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to create DNS records for node {i} ' > + f'({txn.get_error()})') > + > + > +def add_nat_rules(idl, n, routers): > + """Add NAT rules to routers. > + > + Creates SNAT, DNAT, and DNAT_AND_SNAT rules to exercise all NAT code > paths. > + """ > + for i in range(n): > + vlog.info(f'Adding NAT rules to router {i}') > + txn = ovs.db.idl.Transaction(idl) > + > + lr = routers.get(f'lr-{i}') > + if not lr: > + die(f'Missing router lr-{i}') > + > + nat_snat = txn.insert(idl.tables['NAT']) > + nat_snat.type = 'snat' > + nat_snat.logical_ip = f'10.{ip_node(i)}.0/24' > + nat_snat.external_ip = f'192.{ip_node(i)}.1' > + lr.addvalue('nat', nat_snat.uuid) > + > + nat_dnat = txn.insert(idl.tables['NAT']) > + nat_dnat.type = 'dnat' > + nat_dnat.logical_ip = f'10.{ip_node(i)}.10' > + nat_dnat.external_ip = f'192.{ip_node(i)}.10' > + nat_dnat.setkey('external_ids', 'service', 'web') > + lr.addvalue('nat', nat_dnat.uuid) > + > + nat_dnat_and_snat = txn.insert(idl.tables['NAT']) > + nat_dnat_and_snat.type = 'dnat_and_snat' > + nat_dnat_and_snat.logical_ip = f'10.{ip_node(i)}.20' > + nat_dnat_and_snat.external_ip = f'192.{ip_node(i)}.20' > + nat_dnat_and_snat.setkey('external_ids', 'service', 'db') > + lr.addvalue('nat', nat_dnat_and_snat.uuid) > + > + nat_snat_v6 = txn.insert(idl.tables['NAT']) > + nat_snat_v6.type = 'snat' > + nat_snat_v6.logical_ip = f'fd00:{ip6_node(i)}::/64' > + nat_snat_v6.external_ip = f'fd01:{ip6_node(i)}::1' > + lr.addvalue('nat', nat_snat_v6.uuid) > + > + nat_dnat_v6 = txn.insert(idl.tables['NAT']) > + nat_dnat_v6.type = 'dnat' > + nat_dnat_v6.logical_ip = f'fd00:{ip6_node(i)}::a' > + nat_dnat_v6.external_ip = f'fd01:{ip6_node(i)}::a' > + nat_dnat_v6.setkey('external_ids', > + 'service', 'web-v6') > + lr.addvalue('nat', nat_dnat_v6.uuid) > + > + nat_ds_v6 = txn.insert(idl.tables['NAT']) > + nat_ds_v6.type = 'dnat_and_snat' > + nat_ds_v6.logical_ip = f'fd00:{ip6_node(i)}::14' > + nat_ds_v6.external_ip = f'fd01:{ip6_node(i)}::14' > + nat_ds_v6.setkey('external_ids', > + 'service', 'db-v6') > + lr.addvalue('nat', nat_ds_v6.uuid) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to add NAT rules for node {i} ' > + f'({txn.get_error()})') > + > + > +def add_static_routes(idl, n, routers): > + """Add static routes to routers.""" > + for i in range(n): > + vlog.info(f'Adding static routes to router {i}') > + txn = ovs.db.idl.Transaction(idl) > + > + lr = routers.get(f'lr-{i}') > + if not lr: > + die(f'Missing router lr-{i}') > + > + route_default = txn.insert( > + idl.tables['Logical_Router_Static_Route']) > + route_default.ip_prefix = '0.0.0.0/0' > + route_default.nexthop = '100.64.0.1' > + route_default.setkey('external_ids', 'type', 'default') > + lr.addvalue('static_routes', route_default.uuid) > + > + route_specific = txn.insert( > + idl.tables['Logical_Router_Static_Route']) > + route_specific.ip_prefix = f'172.{ip_node(i)}.0/24' > + route_specific.nexthop = f'10.{ip_node(i)}.254' > + route_specific.setkey('external_ids', 'type', 'specific') > + lr.addvalue('static_routes', route_specific.uuid) > + > + route_discard = txn.insert( > + idl.tables['Logical_Router_Static_Route']) > + route_discard.ip_prefix = '192.0.2.0/24' > + route_discard.nexthop = 'discard' > + route_discard.setkey('external_ids', 'type', 'blackhole') > + lr.addvalue('static_routes', route_discard.uuid) > + > + route_default_v6 = txn.insert( > + idl.tables['Logical_Router_Static_Route']) > + route_default_v6.ip_prefix = '::/0' > + route_default_v6.nexthop = 'fd01::1' > + route_default_v6.setkey('external_ids', 'type', > + 'default-v6') > + lr.addvalue('static_routes', route_default_v6.uuid) > + > + route_specific_v6 = txn.insert( > + idl.tables['Logical_Router_Static_Route']) > + route_specific_v6.ip_prefix = ( > + f'fd02:{ip6_node(i)}::/64') > + route_specific_v6.nexthop = ( > + f'fd00:{ip6_node(i)}::fe') > + route_specific_v6.setkey('external_ids', 'type', > + 'specific-v6') > + lr.addvalue('static_routes', > + route_specific_v6.uuid) > + > + route_discard_v6 = txn.insert( > + idl.tables['Logical_Router_Static_Route']) > + route_discard_v6.ip_prefix = '2001:db8::/32' > + route_discard_v6.nexthop = 'discard' > + route_discard_v6.setkey('external_ids', 'type', > + 'blackhole-v6') > + lr.addvalue('static_routes', > + route_discard_v6.uuid) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to add static routes for node {i} ' > + f'({txn.get_error()})') > + > + > +def add_routing_policies(idl, n, routers): > + """Add routing policies to routers.""" > + for i in range(n): > + vlog.info(f'Adding routing policies to router {i}') > + txn = ovs.db.idl.Transaction(idl) > + > + lr = routers.get(f'lr-{i}') > + if not lr: > + die(f'Missing router lr-{i}') > + > + policy_reroute = txn.insert(idl.tables['Logical_Router_Policy']) > + policy_reroute.priority = 100 > + policy_reroute.match = f'ip4.src == 10.{ip_node(i)}.0/24' > + policy_reroute.action = 'reroute' > + policy_reroute.nexthops = [f'10.{ip_node((i + 1) % n)}.1'] > + policy_reroute.setkey('external_ids', 'policy', > + 'traffic-engineering') > + lr.addvalue('policies', policy_reroute.uuid) > + > + policy_allow = txn.insert(idl.tables['Logical_Router_Policy']) > + policy_allow.priority = 50 > + policy_allow.match = 'ip4.dst == $trusted_networks_v4' > + policy_allow.action = 'allow' > + lr.addvalue('policies', policy_allow.uuid) > + > + policy_reroute_v6 = txn.insert( > + idl.tables['Logical_Router_Policy']) > + policy_reroute_v6.priority = 100 > + policy_reroute_v6.match = ( > + f'ip6.src == fd00:{ip6_node(i)}::/64') > + policy_reroute_v6.action = 'reroute' > + policy_reroute_v6.nexthops = [ > + f'fd00:{ip6_node((i + 1) % n)}::1'] > + policy_reroute_v6.setkey( > + 'external_ids', 'policy', > + 'traffic-engineering-v6') > + lr.addvalue('policies', > + policy_reroute_v6.uuid) > + > + policy_allow_v6 = txn.insert( > + idl.tables['Logical_Router_Policy']) > + policy_allow_v6.priority = 50 > + policy_allow_v6.match = ( > + 'ip6.dst == $trusted_networks_v6') > + policy_allow_v6.action = 'allow' > + lr.addvalue('policies', policy_allow_v6.uuid) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to add routing policies for node {i} ' > + f'({txn.get_error()})') > + > + > +def _add_ports(txn, idl, s, i, ports_per_switch): > + for p in range(ports_per_switch): > + lsp = txn.insert(idl.tables['Logical_Switch_Port']) > + lsp.name = f'lsp-{i}-{p}' > + # Offset by 10 to match the IP's +10 offset; wraps at 256 but > + # the port cap (244) means this only matters at the boundary. > + mac_byte = (p + 10) % 256 > + mac = (f'00:00:{i >> 8:02x}:{i & 0xff:02x}' > + f':{p:02x}:{mac_byte:02x}') > + ip = f'10.{ip_node(i)}.{10 + p}' > + ip6 = f'fd00:{ip6_node(i)}::{10 + p:x}' > + lsp.addresses = [f'{mac} {ip} {ip6}'] > + lsp.port_security = [f'{mac} {ip} {ip6}'] > + lsp.setkey('external_ids', 'vm-id', f'vm-{i}-{p}') > + > + if p % 3 == 0: > + lsp.setkey('external_ids', 'tier', 'web') > + elif p % 3 == 1: > + lsp.setkey('external_ids', 'tier', 'app') > + else: > + lsp.setkey('external_ids', 'tier', 'db') > + > + s.addvalue('ports', lsp.uuid) > + > + > +def _add_node(txn, idl, i, chassis, cluster_rtr, join_sw, lbg, > + ports_per_switch): > + gwr = txn.insert(idl.tables['Logical_Router']) > + gwr.name = f'lr-{i}' > + gwr.addvalue('load_balancer_group', lbg.uuid) > + gwr.setkey('options', 'chassis', chassis) > + > + gwr2join = txn.insert(idl.tables['Logical_Router_Port']) > + gwr2join.name = f'lr2j-{i}' > + gwr2join.mac = f'00:00:01:{i >> 8:02x}:{i & 0xff:02x}:01' > + gwr2join.networks = [f'100.64.{ip_node(i)}/16', > + f'fd01:{ip6_node(i)}::1/48'] > + gwr.addvalue('ports', gwr2join.uuid) > + > + join2gwr = txn.insert(idl.tables['Logical_Switch_Port']) > + join2gwr.name = f'j2lr-{i}' > + join2gwr.type = 'router' > + join2gwr.addresses = ['router'] > + join2gwr.setkey('options', 'router-port', gwr2join.name) > + join_sw.addvalue('ports', join2gwr.uuid) > + > + s = txn.insert(idl.tables['Logical_Switch']) > + s.name = f'ls-{i}' > + s.addvalue('load_balancer_group', lbg.uuid) > + s.setkey('other_config', 'subnet', f'10.{ip_node(i)}.0/24') > + s.setkey('other_config', 'mcast_snoop', 'true') > + > + cluster2s = txn.insert(idl.tables['Logical_Router_Port']) > + cluster2s.name = f'c2s-{i}' > + cluster2s.mac = '00:00:00:00:00:01' > + cluster2s.networks = [f'10.{ip_node(i)}.1/24', > + f'fd00:{ip6_node(i)}::1/64'] > + cluster_rtr.addvalue('ports', cluster2s.uuid) > + > + gw_chassis = txn.insert(idl.tables['Gateway_Chassis']) > + gw_chassis.name = f'{cluster2s.name}-{chassis}' > + gw_chassis.chassis_name = chassis > + gw_chassis.priority = 1 > + cluster2s.addvalue('gateway_chassis', gw_chassis.uuid) > + > + s2cluster = txn.insert(idl.tables['Logical_Switch_Port']) > + s2cluster.name = f's2c-{i}' > + s2cluster.type = 'router' > + s2cluster.addresses = ['router'] > + s2cluster.setkey('options', 'router-port', cluster2s.name) > + s.addvalue('ports', s2cluster.uuid) > + > + _add_ports(txn, idl, s, i, ports_per_switch) > + > + > +def create_topology(idl, n, ports_per_switch, batch_size): > + """Create the basic topology with routers, switches, and ports.""" > + vlog.info('Creating topology') > + txn = ovs.db.idl.Transaction(idl) > + lbg = txn.insert(idl.tables['Load_Balancer_Group']) > + lbg.name = 'lbg' > + > + vlog.info('Adding join switch') > + join_sw = txn.insert(idl.tables['Logical_Switch']) > + join_sw.name = 'join' > + > + cluster_rtr = txn.insert(idl.tables['Logical_Router']) > + cluster_rtr.name = 'cluster' > + > + rcj = txn.insert(idl.tables['Logical_Router_Port']) > + rcj.name = 'rcj' > + rcj.mac = '00:00:00:00:00:01' > + rcj.networks = ['100.64.0.1/16', 'fd01::1/48'] > + cluster_rtr.addvalue('ports', rcj.uuid) > + > + sjc = txn.insert(idl.tables['Logical_Switch_Port']) > + sjc.name = 'sjc' > + sjc.type = 'router' > + sjc.addresses = ['router'] > + sjc.setkey('options', 'router-port', 'rcj') > + join_sw.addvalue('ports', sjc.uuid) > + > + for i in range(n): > + vlog.info(f'Provisioning node {i}') > + chassis = f'chassis-{i // batch_size + 1}' > + _add_node(txn, idl, i, chassis, cluster_rtr, join_sw, lbg, > + ports_per_switch) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to create topology ({txn.get_error()})') > + > + > +def assign_ports_to_groups(idl): > + """Assign logical switch ports to port groups based on tier.""" > + vlog.info('Assigning ports to port groups') > + > + tier_ports = {'web': [], 'app': [], 'db': []} > + > + for row in idl.tables['Logical_Switch_Port'].rows.values(): > + tier = row.external_ids.get('tier') > + if tier in tier_ports: > + tier_ports[tier].append(row.uuid) > + > + txn = ovs.db.idl.Transaction(idl) > + port_groups = {row.name: row > + for row in idl.tables['Port_Group'].rows.values()} > + tier_to_group = {'web': 'web_tier', 'app': 'app_tier', 'db': > 'db_tier'} > + for tier, group_name in tier_to_group.items(): > + pg = port_groups.get(group_name) > + if not pg: > + die(f'Missing port group {group_name}') > + for port_uuid in tier_ports[tier]: > + pg.addvalue('ports', port_uuid) > + > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to assign ports to groups ({txn.get_error()})') > + > + > +def add_explicit_lbs(idl, n, n_vips, n_backends, routers, switches): > + """Add explicit (non-templated) load balancers. > + > + Uses explicit (non-templated) LBs to maximize memory usage for > + regression testing. > + """ > + for i in range(n): > + lr = routers.get(f'lr-{i}') > + ls = switches.get(f'ls-{i}') > + if not lr or not ls: > + die(f'Missing router or switch for node {i}') > + for j in range(n_vips): > + vlog.info(f'Adding LB {j} for node {i}') > + txn = ovs.db.idl.Transaction(idl) > + port = j + 1 > + # Spread VIP index across two octets to stay within 0-249. > + # Unlike ip_node() which uses >>8/&0xff for node addresses, > + # this uses //250/%250 because VIP counts are small and 250 > + # avoids addresses that look like broadcast/network octets. > + # Backend addresses depend only on j (not i), so every node > + # shares the same backend set -- intentional, since we are > + # testing memory, not realistic traffic distribution. > + j1 = (j + 1) // 250 > + j2 = (j + 1) % 250 > + > + lb = txn.insert(idl.tables['Load_Balancer']) > + lb.name = f'lb-{j}-{i}' > + lb.setkey('vips', f'42.42.{ip_node(i)}:{port}', > + ','.join(f'42.{k}.{j1}.{j2}:{port}' > + for k in range(n_backends))) > + lb.setkey( > + 'vips', > + f'[fd42::{ip6_node(i)}:{j:x}]:{port}', > + ','.join(f'[fd42:{k:x}::{j1:x}:{j2:x}]:{port}' > + for k in range(n_backends))) > + lb.protocol = 'tcp' > + lr.addvalue('load_balancer', lb.uuid) > + ls.addvalue('load_balancer', lb.uuid) > + if txn.commit_block() != ovs.db.idl.Transaction.SUCCESS: > + die(f'Failed to add LB ({txn.get_error()})') > + > + > +def run(remote, n, n_vips, n_backends, ports_per_switch, batch_size): > + """Main execution function.""" > + schema_helper = ovs.db.idl.SchemaHelper(SCHEMA) > + schema_helper.register_all() > + idl = ovs.db.idl.Idl(remote, schema_helper, leader_only=False) > + > + while idl.change_seqno == 0 and not idl.run(): > + poller = ovs.poller.Poller() > + idl.wait(poller) > + poller.block() > + > + # Check if database is clean before proceeding > + if (len(idl.tables['Load_Balancer_Group'].rows) > 0 or > + len(idl.tables['Logical_Switch'].rows) > 0 or > + len(idl.tables['Logical_Router'].rows) > 0): > + die('Database is not empty. Please restart the sandbox or clear > the ' > + 'database before running this script.') > + > + create_topology(idl, n, ports_per_switch, batch_size) > + > + # Build lookup dictionaries for O(1) access to switches and routers > + switches = {row.name: row > + for row in idl.tables['Logical_Switch'].rows.values()} > + routers = {row.name: row > + for row in idl.tables['Logical_Router'].rows.values()} > + > + create_address_sets(idl, n) > + create_port_groups(idl) > + assign_ports_to_groups(idl) > + create_dhcp_options(idl, n) > + assign_dhcp_options(idl, n) > + create_dns_records(idl, n, switches) > + add_nat_rules(idl, n, routers) > + add_static_routes(idl, n, routers) > + create_acls(idl) > + for i in range(n): > + add_acls_to_switch(idl, f'ls-{i}', i, switches) > + create_qos_rules(idl, n, switches) > + add_routing_policies(idl, n, routers) > + add_explicit_lbs(idl, n, n_vips, n_backends, routers, switches) > + > + > +def main(): > + """Parse arguments and run the benchmark.""" > + parser = argparse.ArgumentParser( > + description='Create a complex OVN topology with various features' > + ) > + parser.add_argument( > + '-r', '--remote', required=True, help='NB connection string' > + ) > + parser.add_argument( > + '-n', '--nodes', type=int, required=True, help='Number of nodes' > + ) > + parser.add_argument( > + '-p', > + '--ports-per-switch', > + type=int, > + default=9, > + help='Number of logical switch ports per switch (default: 9, ' > + 'provides 3 ports per tier)', > + ) > + parser.add_argument( > + '-v', '--vips', type=int, default=5, > + help='Number of LB VIPs per node (default: 5)' > + ) > + parser.add_argument( > + '-B', > + '--backends', > + type=int, > + default=5, > + help='Number backends per VIP (default: 5)', > + ) > + parser.add_argument( > + '-b', > + '--batch-size', > + type=int, > + default=0, > + help='Nodes per chassis (default: n/10)', > + ) > + parser.add_argument( > + '-d', '--debug', > + action='store_true', > + help='Enable debug output (show info messages)', > + ) > + args = parser.parse_args() > + > + if args.batch_size <= 0: > + args.batch_size = max(1, args.nodes // 10) > + > + # Node index is split across two IP octets via ip_node(), so the > + # maximum is 2^16 - 1. > + if args.nodes > 65535: > + sys.stderr.write('Error: maximum supported node count is 65535\n') > + sys.exit(1) > + > + # Backend addresses use 42.{k}.j1.j2; k must fit in one octet. > + if args.backends > 255: > + sys.stderr.write('Error: maximum supported backends is 255\n') > + sys.exit(1) > + > + # VIP index is split via (j+1)//250 and (j+1)%250; the quotient > + # must fit in one octet, so max VIPs = 250 * 256 - 1 = 63999. > + if args.vips > 63999: > + sys.stderr.write('Error: maximum supported VIPs is 63999\n') > + sys.exit(1) > + > + # Port IPs start at 10.x.y.10 (see _add_ports), so the highest > + # address is 10.x.y.(10 + ports_per_switch - 1). Cap at 245 to > + # keep that <= 254 (valid IPv4 octet). > + if args.ports_per_switch > 245: > + sys.stderr.write( > + 'Error: maximum supported ports per switch is 245\n') > + sys.exit(1) > + > + if args.debug: > + vlog.set_levels_from_string('console:info') > + > + # Print configuration summary > + sys.stderr.write('\n=== OVN Benchmark Configuration ===\n') > + sys.stderr.write(f'Nodes: {args.nodes}' > + f' ({args.nodes} routers + {args.nodes} switches)\n') > + sys.stderr.write(f'Ports per switch: > {args.ports_per_switch} ' > + f'({args.ports_per_switch * args.nodes} total > ports)\n') > + sys.stderr.write(f'Load balancer VIPs per node: {args.vips} ' > + f'({args.vips * args.nodes} total VIPs)\n') > + sys.stderr.write(f'Backends per VIP: {args.backends}\n') > + sys.stderr.write(f'Total load balancers: ' > + f'{args.nodes * args.vips}\n') > + n_chassis = ((args.nodes + args.batch_size - 1) > + // args.batch_size) > + sys.stderr.write(f'Nodes per chassis (batch): ' > + f'{args.batch_size} ' > + f'({n_chassis} chassis)\n') > + sys.stderr.write(f'Debug logging: ' > + f'{"enabled" if args.debug else "disabled"}\n') > + sys.stderr.write('===================================\n\n') > + > + run(args.remote, args.nodes, args.vips, args.backends, > + args.ports_per_switch, args.batch_size) > + > + > +if __name__ == '__main__': > + try: > + main() > + except error.Error as e: > + sys.stderr.write(f'{e}\n') > + sys.exit(1) > diff --git a/tutorial/ovn-benchmark.sh b/tutorial/ovn-benchmark.sh > new file mode 100755 > index 000000000..4569813ae > --- /dev/null > +++ b/tutorial/ovn-benchmark.sh > @@ -0,0 +1,467 @@ > +#!/bin/bash > + > +DEFAULT_NODES=200 > + > +PROCESS_NAME=() > +FILE_NAME="" > +NODES="" > +PROCESS_PIDS=() > +FINAL_PEAK_KB=() > +FINAL_PEAK_MB=() > +DEBUG=false > +BATCH_SIZE="" > +VALGRIND_MODE=false > +VALGRIND_PIDS=() > +MASSIF_FILES=() > +BENCHMARK_TMPDIR="" > + > +on_interrupt() { > + echo "" > + if [ "$VALGRIND_MODE" = true ]; then > + echo "Exiting benchmark. Restarting tracked sandbox processes..." > + fi > + exit 1 > +} > +trap on_interrupt INT > + > +# In valgrind mode, the original daemons are killed and replaced with > +# valgrind-wrapped copies. This cleanup ensures they are restored on > +# both normal exit and Ctrl+C so the sandbox is not left broken. > +cleanup() { > + # Stop any valgrind processes still running. > + for vpid in "${VALGRIND_PIDS[@]}"; do > + kill "$vpid" 2>/dev/null > + done > + > + if [ -n "$BENCHMARK_TMPDIR" ]; then > + deadline=$(($(date +%s) + 30)) > + for vpid in "${VALGRIND_PIDS[@]}"; do > + while [ -e "/proc/$vpid" ]; do > + if [ "$(date +%s)" -ge "$deadline" ]; then > + echo "Warning: valgrind PID $vpid" \ > + "did not exit; killing" > + kill -9 "$vpid" 2>/dev/null > + break > + fi > + sleep 0.1 > + done > + done > + > + # Restart each daemon from the cmdline/cwd saved before we > + # killed it, but only if it is not already running. > + for i in "${!PROCESS_NAME[@]}"; do > + pn="${PROCESS_NAME[$i]}" > + if [ -f "$BENCHMARK_TMPDIR/cmdline.$pn" ] \ > + && ! pgrep -x "$pn" >/dev/null 2>&1; then > + restart_cwd=$(cat "$BENCHMARK_TMPDIR/cwd.$pn") > + mapfile -d '' restart_args < > "$BENCHMARK_TMPDIR/cmdline.$pn" > + # Drop trailing empty element from /proc/*/cmdline's > final NUL. > + if [ ${#restart_args[@]} -gt 0 ] \ > + && [ -z "${restart_args[-1]}" ]; then > + unset 'restart_args[-1]' > + fi > + (cd "$restart_cwd" && "${restart_args[@]}" >/dev/null > 2>&1) > + fi > + done > + > + rm -rf "$BENCHMARK_TMPDIR" > + fi > +} > +trap cleanup EXIT > + > +while [[ $# -gt 0 ]]; do > + case "$1" in > + -h|--help|--usage) > + echo "Usage: $0 [OPTIONS] [NODES] [PROCESS...]" > + echo "" > + echo "Arguments:" > + echo " NODES Number of nodes to create" \ > + "(default: $DEFAULT_NODES)" > + echo " PROCESS Process(es) to track:" \ > + "ovn-northd, ovn-controller" > + echo " (default: both)" > + echo "" > + echo "Options:" > + echo " -f, --file FILE Load NB database from file" > + echo " -b, --batch-size N Nodes per chassis" \ > + "(default: NODES/10)" > + echo " -v, --valgrind Track heap with Valgrind" \ > + "Massif (slower, accurate)" > + echo " -d, --debug Enable debug output" > + echo " -h, --help Show this help message" > + echo "" > + echo "Memory tracking:" > + echo " Default: peak virtual memory (VmPeak) from" > + echo " /proc/<pid>/status." > + echo " --valgrind: peak heap via Valgrind Massif." > + echo " Expect 10-50x slowdown; 500 nodes or fewer" > + echo " recommended." > + echo "" > + echo "Examples:" > + echo " $0 # 200 nodes, track both > processes" > + echo " $0 50 # 50 nodes" > + echo " $0 50 ovn-northd # 50 nodes, track only > ovn-northd" > + echo " $0 --valgrind # Use Valgrind Massif, not > VmPeak" > + echo " $0 --file ovnnb_db.db # Load from file" > + echo " $0 --debug 20 # 20 nodes with debug output" > + echo " $0 50 -b 10 # 50 nodes, 10 per chassis" > + echo "" > + echo "Note: if only tracking one process, # of nodes is > required" + exit 0 > + ;; > + -b|--batch-size) > + if [ -z "$2" ]; then > + echo "Error: $1 requires an argument" > + exit 1 > + fi > + BATCH_SIZE="$2" > + shift 2 > + ;; > + -v|--valgrind) > + VALGRIND_MODE=true > + shift > + ;; > + -d|--debug) > + DEBUG=true > + shift > + ;; > + -f|--file) > + if [ -z "$2" ]; then > + echo "Error: $1 requires an argument" > + exit 1 > + fi > + FILE_NAME="$2" > + shift 2 > + ;; > + -*) > + echo "Unknown option: $1" > + exit 1 > + ;; > + *) > + if [ -z "$NODES" ]; then > + NODES="$1" > + else > + # Normalize process names: accept both "northd" and > + # "ovn-northd". > + case "$1" in > + northd) > + PROCESS_NAME+=("ovn-northd") > + ;; > + controller) > + PROCESS_NAME+=("ovn-controller") > + ;; > + *) > + PROCESS_NAME+=("$1") > + ;; > + esac > + fi > + shift > + ;; > + esac > +done > + > +# Must be run from inside the sandbox (tutorial/sandbox/). > +if [ ! -d "$PWD/sandbox" ]; then > + echo "Error: must be run from inside the OVN sandbox." > + echo "Start one with: make sandbox" > + exit 1 > +fi > + > +# Apply defaults if not set by user. > +NODES=${NODES:-$DEFAULT_NODES} > + > +if [ -z "$BATCH_SIZE" ]; then > + BATCH_SIZE=$((NODES / 10)) > +fi > +if [ "$BATCH_SIZE" -lt 1 ]; then > + BATCH_SIZE=1 > +fi > + > +# Track both processes if not specified. > +if [ ${#PROCESS_NAME[@]} -eq 0 ]; then > + PROCESS_NAME=("ovn-controller" "ovn-northd") > +fi > + > +if [ "$DEBUG" = true ]; then > + echo "Nodes: $NODES" > + echo "Batch size: $BATCH_SIZE" > + echo "Processes: ${PROCESS_NAME[*]}" > + echo "File: ${FILE_NAME:-None}" > +fi > + > +for pn in "${PROCESS_NAME[@]}"; do > + all_pids=$(pgrep -x "$pn") > + if [ -z "$all_pids" ]; then > + echo "Error: Could not find process matching '$pn'" > + exit 1 > + fi > + > + if [ "$(echo "$all_pids" | wc -l)" -gt 1 ]; then > + echo "Error: Multiple $pn processes found" \ > + "(PIDs: $(echo $all_pids | tr '\n' ' '))" > + echo "Kill stale processes or ensure only one sandbox is running." > + exit 1 > + fi > + > + PROCESS_PIDS+=("$all_pids") > +done > + > +if [ "$DEBUG" = true ]; then > + for i in "${!PROCESS_NAME[@]}"; do > + echo "Tracking memory for ${PROCESS_NAME[$i]}" \ > + "(PID: ${PROCESS_PIDS[$i]})" > + done > +fi > + > +if [ "$VALGRIND_MODE" = true ]; then > + if ! command -v valgrind >/dev/null 2>&1; then > + echo "Error: valgrind is not installed or not in PATH" > + exit 1 > + fi > + > + if [ "$NODES" -gt 500 ]; then > + echo "Warning: valgrind mode with $NODES nodes may be very slow." > \ > + "Consider using 500 or fewer nodes for heap profiling." > + fi > + > + BENCHMARK_TMPDIR=$(mktemp -d) > + echo "Restarting processes under Valgrind Massif..." > + echo "(Expect 10-50x slowdown)" > + > + for i in "${!PROCESS_NAME[@]}"; do > + pn="${PROCESS_NAME[$i]}" > + pid="${PROCESS_PIDS[$i]}" > + > + # Read the original command line so we can restart the daemon > later. > + # /proc/<pid>/cmdline is NUL-separated; the trailing NUL produces > an > + # empty final element which must be dropped. > + mapfile -d '' orig_args < /proc/$pid/cmdline > + if [ ${#orig_args[@]} -gt 0 ] \ > + && [ -z "${orig_args[-1]}" ]; then > + unset 'orig_args[-1]' > + fi > + > + proc_cwd=$(readlink /proc/$pid/cwd) > + massif_file="$BENCHMARK_TMPDIR/massif.$pn.out" > + MASSIF_FILES+=("$massif_file") > + > + # Save original cmdline and cwd so we can restart after valgrind. > + cp /proc/$pid/cmdline "$BENCHMARK_TMPDIR/cmdline.$pn" > + echo "$proc_cwd" > "$BENCHMARK_TMPDIR/cwd.$pn" > + > + # Valgrind runs the process in the foreground, so strip args > + # that assume daemonized execution: > + # --detach : can't daemonize under valgrind > + # --no-chdir : only meaningful with --detach > + # --pidfile : valgrind doesn't write one (OVS uses > + # optional_argument so bare --pidfile never > + # takes a separate arg; only --pidfile=FILE) > + # --monitor : would fork a monitor child under valgrind, > + # causing two processes to race on massif output > + filtered_args=() > + for arg in "${orig_args[@]}"; do > + case "$arg" in > + --detach|--no-chdir|--monitor) ;; > + --pidfile|--pid-file) ;; > + --pidfile=*|--pid-file=*) ;; > + *) filtered_args+=("$arg") ;; > + esac > + done > + > + if [ "$DEBUG" = true ]; then > + echo "Stopping $pn (PID $pid)..." > + fi > + > + kill "$pid" > + while [ -e "/proc/$pid" ]; do sleep 0.1; done > + > + if [ "$DEBUG" = true ]; then > + echo "Starting $pn under valgrind:" > + echo " valgrind --tool=massif > --massif-out-file=$massif_file" \ > + "${filtered_args[*]}" > + fi > + > + # exec replaces the subshell with valgrind so $! is valgrind's > + # actual PID and SIGTERM reaches it to finalize massif output. > + # --max-snapshots, --detailed-freq, and --depth reduce the cost > + # of each snapshot, which otherwise scales with the number of > + # live allocations and becomes prohibitive at large node counts. > + (cd "$proc_cwd" && exec valgrind \ > + --tool=massif \ > + --massif-out-file="$massif_file" \ > + --max-snapshots=50 \ > + --detailed-freq=20 \ > + --depth=10 \ > + "${filtered_args[@]}" \ > + >/dev/null 2>&1) & > + VALGRIND_PIDS+=("$!") > + done > + > + # Poll until northd and controller are processing. > + # --wait=hv confirms the full stack is functional. > + # (NB -> northd -> SB -> controller). > + echo "Waiting for processes to reconnect..." > + deadline=$(($(date +%s) + 60)) > + while ! ovn-nbctl --timeout=5 --wait=hv sync 2>/dev/null; do > + if [ "$(date +%s)" -ge "$deadline" ]; then > + echo "Warning: processes did not reconnect within 60 seconds" > + break > + fi > + sleep 1 > + done > +fi > + > +if [ "$DEBUG" = true ]; then > + DEBUG_FLAG="-d" > +else > + DEBUG_FLAG="" > +fi > + > +# %s%2N gives epoch seconds with two fractional digits (hundredths). > +GEN_START=$(date +%s%2N) > + > +# Load database from file or generate with Python script. > +if [ -n "$FILE_NAME" ]; then > + echo "Loading database from file: $FILE_NAME" > + if [ ! -f "$FILE_NAME" ]; then > + echo "Error: File '$FILE_NAME' not found" > + exit 1 > + fi > + if [ "$VALGRIND_MODE" != true ]; then > + echo "Warning: -f mode uses VmPeak, which is a lifetime" \ > + "high-water mark." > + echo "Restart the sandbox before each run for accurate readings." > + fi > + ovsdb-client restore "unix:$PWD/sandbox/nb1.ovsdb" < "$FILE_NAME" > +else > + echo "Generating database with Python script" > + python3 ovn-benchmark.py -n "$NODES" -b "$BATCH_SIZE" \ > + -r "unix:$PWD/sandbox/nb1.ovsdb" $DEBUG_FLAG > + if [ $? -ne 0 ]; then > + echo "Error: Failed to generate database" > + exit 1 > + fi > +fi > + > +# Bind the first port of each switch assigned to chassis-0. > +BIND_COUNT=$((BATCH_SIZE < NODES ? BATCH_SIZE : NODES)) > +for i in $(seq 0 $((BIND_COUNT - 1))); do > + ovs-vsctl add-port br-int lsp-${i}-0 -- \ > + set interface lsp-${i}-0 \ > + external_ids:iface-id=lsp-${i}-0 > +done > + > +GEN_END=$(date +%s%2N) > + > +# Time OVN processing separately from DB generation. > +PROC_START=$(date +%s%2N) > +ovn-nbctl --wait=hv sync > should this have a timeout like the valgrind reconnect loop (which uses --timeout=5)? If a chassis never comes up will the benchmark hang? > +PROC_END=$(date +%s%2N) > + > +# Compact before measuring memory. > +COMPACT_START=$(date +%s%2N) > +ovs-appctl -t "$PWD/sandbox/nb1" ovsdb-server/compact > +ovs-appctl -t "$PWD/sandbox/sb1" ovsdb-server/compact > +COMPACT_END=$(date +%s%2N) > + > +GEN_ELAPSED=$((GEN_END - GEN_START)) > +PROC_ELAPSED=$((PROC_END - PROC_START)) > +COMPACT_ELAPSED=$((COMPACT_END - COMPACT_START)) > +TOTAL_ELAPSED=$((COMPACT_END - GEN_START)) > + > +if [ "$VALGRIND_MODE" = true ]; then > + if [ "$DEBUG" = true ]; then > + echo "Benchmark complete. Stopping valgrind and collecting > results..." > + fi > + > + # Signal all valgrind processes to stop. Each valgrind instance > + # propagates the signal to its child OVN process, waits for it to > + # exit, then writes the massif output file and exits itself. > + for i in "${!PROCESS_NAME[@]}"; do > + pn="${PROCESS_NAME[$i]}" > + vpid="${VALGRIND_PIDS[$i]}" > + if [ "$DEBUG" = true ]; then > + echo "Stopping valgrind for $pn (PID $vpid)..." > + fi > + kill "$vpid" 2>/dev/null > + done > + > + deadline=$(($(date +%s) + 60)) > + for i in "${!PROCESS_NAME[@]}"; do > + pn="${PROCESS_NAME[$i]}" > + vpid="${VALGRIND_PIDS[$i]}" > + if [ "$DEBUG" = true ]; then > + echo "Waiting for valgrind to write massif output for $pn..." > + fi > + while [ -e "/proc/$vpid" ]; do > + if [ "$(date +%s)" -ge "$deadline" ]; then > + echo "Warning: valgrind for $pn did not exit" > + break > + fi > + sleep 0.1 > + done > + if [ "$DEBUG" = true ]; then > + echo "Massif output for $pn ready." > + fi > + done > + > + for i in "${!PROCESS_NAME[@]}"; do > + pn="${PROCESS_NAME[$i]}" > + massif_file="${MASSIF_FILES[$i]}" > + if [ "$DEBUG" = true ]; then > + echo "Parsing massif output for $pn..." > + fi > + if [ ! -s "$massif_file" ]; then > + echo "Warning: massif output missing for $pn;" \ > + "did valgrind exit cleanly?" > + FINAL_PEAK_KB[$i]=0 > + else > + # Sum heap + allocator overhead per snapshot; report the peak. > + FINAL_PEAK_KB[$i]=$(awk -F= ' > + /^mem_heap_B=/{h=$2} > + /^mem_heap_extra_B=/{e=$2; t=h+e; if(t>peak) peak=t} > + END{print int(peak/1024)} > + ' "$massif_file") > + fi > + FINAL_PEAK_MB[$i]=$((FINAL_PEAK_KB[$i] / 1024)) > + done > + > + # Daemons are restarted by cleanup() on EXIT. > +else > + for i in "${!PROCESS_NAME[@]}"; do > + pid=${PROCESS_PIDS[$i]} > + FINAL_PEAK_KB[$i]=$(awk '/^VmPeak:/{print $2}' \ > + /proc/$pid/status 2>/dev/null) > + if [ -z "${FINAL_PEAK_KB[$i]}" ]; then > + echo "Warning: ${PROCESS_NAME[$i]} (PID $pid)" \ > + "is no longer running" > + FINAL_PEAK_KB[$i]=0 > + fi > + FINAL_PEAK_MB[$i]=$((FINAL_PEAK_KB[$i] / 1024)) > + done > +fi > + > +echo "" > +echo "=== Benchmark Results ===" > +printf "Total time: %d.%02d seconds\n" \ > + $((TOTAL_ELAPSED / 100)) $((TOTAL_ELAPSED % 100)) > +printf " (topology gen %d.%02ds + OVN sync %d.%02ds + \ > +db compact %d.%02ds)\n" \ > + $((GEN_ELAPSED / 100)) $((GEN_ELAPSED % 100)) \ > + $((PROC_ELAPSED / 100)) $((PROC_ELAPSED % 100)) \ > + $((COMPACT_ELAPSED / 100)) $((COMPACT_ELAPSED % 100)) > +echo "" > +if [ "$VALGRIND_MODE" = true ]; then > + echo "Peak heap memory (Valgrind Massif):" > +else > + echo "Peak virtual memory (VmPeak):" > +fi > +for i in "${!PROCESS_NAME[@]}"; do > + printf " %-15s %d.%d MB (%d KB)\n" \ > + "${PROCESS_NAME[$i]}:" \ > + "$((FINAL_PEAK_KB[$i] / 1024))" \ > + "$((FINAL_PEAK_KB[$i] % 1024 * 10 / 1024))" \ > + "${FINAL_PEAK_KB[$i]}" > +done > +echo "=========================" > +echo "" > -- > 2.55.0 > > _______________________________________________ > dev mailing list > [email protected] > https://mail.openvswitch.org/mailman/listinfo/ovs-dev > > Jacob Tanenbaum _______________________________________________ dev mailing list [email protected] https://mail.openvswitch.org/mailman/listinfo/ovs-dev
