gpui: Add seeded benchmarks and heap allocation metrics (#64843)

Two additions to GPUI's benchmark harness, split out of #64842 so the
randomized element tree harness can build on them.

## Seeded `#[gpui::bench]` benchmarks

A benchmark that takes a `StdRng` parameter runs once per seed, and each
seed is its own Criterion benchmark named `<input>/seed-<n>` (or
`<function>/seed-<n>` without inputs). Seeds come from
`calculate_seeds`, moved out of `test.rs` into `seeds.rs` so it compiles
under `bench-support` and is shared with `#[gpui::test]`. `seed = N`,
`seeds(...)`, `iterations = N`, `SEED` and `ITERATIONS` therefore mean
the same in tests and benchmarks, and one `SEED` reproduces either. The
RNG is rebuilt on every Criterion routine call, so warm-up and every
sample measure the same workload. The generated code seeds it through
`gpui::private::rand`, so a bench crate needs `rand` only to name
`StdRng` in its signature.

```rust
#[gpui::bench(inputs = families(), input_name = "tree", iterations = 6)]
fn full_refresh(family: &Family, mut rng: StdRng, cx: &mut BenchAppContext) { ... }
```

The macro now also checks the benchmark's parameters, and reports a
wrong parameter count, a `&mut StdRng`, or `seed`/`seeds`/`iterations`
without a `StdRng` parameter as clear errors.

**Behavior change in `gpui::test`:** with `SEED` set and `iterations >
1`, `calculate_seeds` returned `SEED` and then `SEED..SEED+iterations`,
so `SEED=7 ITERATIONS=2` ran 7, 7, 8. Criterion rejects the duplicate
benchmark ID, and the function's comment already described
`SEED..SEED+iterations`, so it now runs 7, 8.

## Heap allocation metrics in `bench_metrics`

- `CountingAllocator<A = System>` wraps any global allocator and counts
allocations, allocated bytes and freed bytes in sharded relaxed atomics,
with no backtraces. Each thread also counts its own allocations in plain
thread-locals.
- `AllocationCounter` measures allocations or allocated bytes per
iteration over a `MetricScope`, like `HardwareCounter`. When the
allocator is installed, `allocations`, `foreground allocations` and
`allocated bytes` are default secondaries.
`BENCH_MEASUREMENT=allocations` or `foreground-allocations` makes one of
them the primary. Without the allocator they're skipped with a one-time
note, like unavailable hardware counters.
- `foreground allocations` leaves out background workers' timing noise.
It can still vary between samples when the workload does, for example
when frames coalesce or caches grow in early samples, so a CI gate on it
needs a small tolerance.
- `allocation_stats()` exposes the totals and `live_bytes()` for heap
held around setup and teardown.
- `gpui::bench_main!` installs `CountingAllocator` over the system
allocator. `gpui::bench_main!(allocator = mimalloc::MiMalloc; benches)`
wraps another allocator instead, and `allocator = none;` installs
nothing.

Example, `inspector_render` smallest case: `allocations per iteration:
median 558`, `allocated bytes per iteration: median 390 K`.

## Notes for reviewers

- The counting allocator is installed in every GPUI bench binary. It
adds a few relaxed atomic increments per allocation, so wall-time
baselines saved before this PR shift slightly.
- Placements in GPUI's element arena aren't heap allocations, so they
aren't counted once the arena is warm. Arena bytes per iteration could
be a GPUI secondary in a follow-up.

## Validation

- Throwaway seeded bench (inputs `a`/`b`, `iterations = 2`,
`seeds(40)`): correct seed order by default and with `SEED=7`, `SEED=7
ITERATIONS=2`, `ITERATIONS=3`; every routine call drew the same first
value for a seed; the allocations primary reported exactly 1 allocation
/ 64 bytes per iteration for `vec![0u8; 64]`. A second throwaway bench
imported `StdRng` through `gpui::private::rand` with no `rand`
dependency of its own and ran `probe/seed-0` and `probe/seed-1`.
- `inspector_render` ran with the default allocator and with `allocator
= none;` (which prints the note and reports no allocation lines), and
built with `allocator = std::alloc::System;`.
- New unit tests for seed order, process-wide allocation counting, and
exact calling-thread counts that exclude other threads; `bench_metrics`
and `bench_context` tests pass.
- `./script/clippy` and `cargo fmt` clean; `--no-run` builds of
`gpui_platform` and `benchmarks` benches.

Release Notes:

- [GPUI] Added seeded `#[gpui::bench]` benchmarks that take a `StdRng`,
and heap allocation counts in benchmark reports
2440236055Anthony Eid committed on 9/28/2026, 1:02:30 PM· committed by GitHubparent93c87da
9 files changedLine totals unavailable