# compiler_rt: optimize udivmod large-divisor case with trial quotient · gitcafe/zig

[View on GitCafe](https://git.cafe/gitcafe/zig/commit/4fa465fc8f3b6140e635186b4f0e5acae924adcc)

Repository: [gitcafe/zig](https://git.cafe/gitcafe/zig)

Visibility: public

Requested revision: 4fa465fc8f3b6140e635186b4f0e5acae924adcc

Requested commit: 4fa465fc8f3b6140e635186b4f0e5acae924adcc

Commit: 4fa465fc8f3b6140e635186b4f0e5acae924adcc

Tree: f9ccc7a2e677653476ab704abeb4408f40e9e7fb

Author: Koko Bhadra

Committer: Andrew Kelley

## Message

```
compiler_rt: optimize udivmod large-divisor case with trial quotient

Replace the O(n) shift-subtract loop with a constant-time trial
quotient approach (Knuth Algorithm D, TAOCP Vol 2 Section 4.3.1).

The old code iterates clz(b_hi)-clz(a_hi)+1 times (up to 64
iterations of 128-bit arithmetic). The new code uses a single
divwide call to get a trial quotient, then verifies with two
native-width widening multiplies.

Benchmark (Apple M1, ReleaseFast):
- Large divisor, large shift: 87ns -> 7.5ns (11.5x faster)
- Small divisor / uniform: unchanged

```

## Parents

- [15f0af09d0ba8e8a6a7b44702913ca309afe2296](https://git.cafe/gitcafe/zig/commit/15f0af09d0ba8e8a6a7b44702913ca309afe2296?format=markdown)

[Source at this commit](https://git.cafe/gitcafe/zig/tree/4fa465fc8f3b6140e635186b4f0e5acae924adcc?format=markdown)
