mmodzelewski opened a new pull request, #4318: URL: https://github.com/apache/iggy/pull/4318
The server did not raise its open-file limit. If a storage open failed with EMFILE or ENFILE, the shard kept serving and failed each later write. Accept loops retried at once and spun the shard at full CPU, because the kernel keeps the connection in the backlog. At startup, the server now raises the soft RLIMIT_NOFILE to the hard limit. On macOS, it uses OPEN_MAX if the hard limit is higher. If a storage open fails with EMFILE or ENFILE, the process now stops with exit status 4. Every later open fails too, so a supervisor restart is the repair. fatal() moves to server_common and writes to stderr, because exit stops the log appenders before they flush. Accept loops wait one second after EMFILE or ENFILE. They do not stop the process, because a client with many sockets can then stop the node. The new [metadata] partitions_max limits the partitions of a node. CreateTopic and CreatePartitions past the limit fail before consensus with PartitionsLimitReached (2022). The limit is soft, because concurrent creates can go over it. 0 means no limit. Shard 0 logs process and host usage, with open descriptors, every logging.sysinfo_print_interval (10 s by default, 0 disables it). Partition groups log the restored-view line at debug, because one INFO line per partition flooded the boot log. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
