# std.crypto.chacha: support larger vectors on AVX2 and AVX512 targets \(\#15809\) · gitcafe/zig

[View on GitCafe](https://git.cafe/gitcafe/zig/commit/5af89b3dccf7ee375f68e9cd3ee4980fef89e38f)

Repository: [gitcafe/zig](https://git.cafe/gitcafe/zig)

Visibility: public

Requested revision: 5af89b3dccf7ee375f68e9cd3ee4980fef89e38f

Requested commit: 5af89b3dccf7ee375f68e9cd3ee4980fef89e38f

Commit: 5af89b3dccf7ee375f68e9cd3ee4980fef89e38f

Tree: 5262e6f92921170a2df1309f97d15669a7ffc614

Author: Frank Denis

Committer: GitHub

## Message

```
std.crypto.chacha: support larger vectors on AVX2 and AVX512 targets (#15809)

* std.crypto.chacha: support larger vectors on AVX2 and AVX512 targets

Ryzen 7 7700, ChaCha20/8 stream, long outputs:

Generic: 3268 MiB/s
AVX2   : 6023 MiB/s
AVX512 : 8086 MiB/s

Bump the rand.chacha buffer a tiny bit to take advantage of this.
More than 8 blocks doesn't seem to make any measurable difference.

ChaChaPoly also gets a small performance boost from this, albeit
Poly1305 remains the bottleneck.

Generic:  707 MiB/s
AVX2   :  981 MiB/s
AVX512 : 1202 MiB/s

aarch64 appears to generally benefit from 4-way vectorization.

Verified on Apple Silicon, but also on a Cortex A72.
```

## Parents

- [eef92753c7cf677191adc40a7cdf7561ceb43bdb](https://git.cafe/gitcafe/zig/commit/eef92753c7cf677191adc40a7cdf7561ceb43bdb?format=markdown)

[Source at this commit](https://git.cafe/gitcafe/zig/tree/5af89b3dccf7ee375f68e9cd3ee4980fef89e38f?format=markdown)
