Faster BLAKE3 implementation (#25574)

This is a rewrite of the BLAKE3 implementation, with vectorization.

On Apple Silicon, the new implementation is about twice as fast as the previous one.

With AVX2, it is more than 4 times faster.

With AVX512, it is more than 7.5x faster than the previous implementation (from 678 MB/s to 5086 MB/s).
6669885aa2Frank Denis committed on 10/15/2025, 12:03:56 PM· committed by GitHubparent70c21fd
1 file changedLine totals unavailable