> Modern GPUs and AI accelerators are increasingly connected through scale-up > fabrics such as AMD xGMI today and emerging UALink systems [4]. Linux lacks > common, vendor-neutral infrastructure for reporting which accelerators are > directly connected, through which ports, and in what state.
Sorry for jumping in late to this discussion. We have hardware that uses UALoE (https://www.amd.com/en/products/rackscale-solutions/helios.html). The suggestion from Jason is sensible: a new subsystem for fabrics and reuse ethernet stuff from networking subsystem where possible. We would like to continue discussing this - Alejandro and I will be attending Plumbers virtually. Thanks, Alex
