Are the numbers in the H100 PCIE vs SXM table swapped for rows 3 onwards? It looks to me like the PCI is showing higher GiB/s numbers, which is counter to expectations.
Or am I misunderstanding those benchmarks?
The point of this post isn’t the linear transformer algorithm. They’re surveying a variety of Linear transformers and showing a general form in order to talk at large about their performance characteristics.
Actually, in the last few months there has been an absolute ton of feature work. A very sophisticated generational profile system was introduced, and the entire framework is being split into separate packages to enable more stable versioning. This is all with the goal of reducing the bus-factor from being just Henrik.
Once you EOL a product, you no longer have to support it. Not the same case at all as making changes to a core product and then continuing to support it.