gpui: Add bench_metrics for multi-metric Criterion benchmarks (#64753)
GPUI benchmarks measured only wall time, which is noisy enough on shared
CI runners that a regression must be large before Criterion flags it.
This PR adds a `bench_metrics` crate that records several metrics per
Criterion sample and reports them alongside Criterion's own analysis,
and wires GPUI's benchmark harness to it. Nothing in `bench_metrics`
depends on GPUI, so any Criterion benchmark in the workspace can adopt
it.
**What a benchmark now prints** (Linux, hardware counters available):
```
closed_inspector_grid/case/100-nodes-0-cycles
time: [273.99 µs 274.78 µs 276.04 µs]
GPUI bench report (all observed iterations): closed_inspector_grid/case/100-nodes-0-cycles
instructions per iteration: median 5.791 M instructions (min 5.788, max 5.880, samples 22)
foreground instructions per iteration: median 5.791 M instructions (...)
cycles per iteration: median 1.453 M cycles (...)
branch misses per iteration: median 1.856 K branch misses (...)
cache misses per iteration: median 123.008 cache misses (...)
foreground context switches per iteration: median 0.000 context switches (...)
page faults per iteration: median 0.000 page faults (...)
IPC: median 3.996 (min 3.349, max 4.063, samples 22)
foreground share of instructions: median 1.000 (...)
```
Criterion analyzes one metric (the *primary*; wall time by default) and
the rest are *secondaries* taken over the same iterations.
`BENCH_MEASUREMENT=instructions` makes process-wide retired instructions
the primary, and `BENCH_MEASUREMENT=foreground-instructions` the
benchmark thread's alone; both are near-deterministic, so a ~0.01%
interval is typical and small regressions are detectable, with wall time
demoted to the report. `BENCH_MEASUREMENT=wall-time` disables all
counters. Counters that can't be opened (no `CAP_PERFMON`, Docker
seccomp, VMs without a PMU, Intel Macs) are skipped with one note, so
`cargo bench` never fails by default; the instruction modes fail fast so
CI can't silently measure the wrong thing.
**Sources.** `HardwareCounter` is platform-neutral over per-OS backends.
On Linux it uses `perf_event_open`: counters are opened once per process
with `inherit` so every later thread is covered, read as deltas per
sample (~150 ns per counter), with hybrid-CPU encodings parsed from
sysfs and multiplexing scaled with a one-time warning. On Apple Silicon
it reads `proc_pid_rusage` and `thread_selfcounts`, which are
unprivileged, for instructions and cycles; branch and cache events need
`kperf` and stay Linux-only. `ResourceCounter` reads `getrusage` for
voluntary context switches on the foreground thread (the waits
instructions can't see) and minor page faults, with no privileges.
Foreground-thread instructions are counted separately from the process
total, so work shifting between GPUI's foreground and background threads
is visible even when the total is flat; with a headless renderer the gap
is driver-thread submission work (about 7% of a small Metal frame).
**Design.** `BenchMeasurement` type-erases the primary so
`#[gpui::bench]` functions and `BenchAppContext` need no type parameter;
`with_secondary`/`with_ratio` add metrics; `bench_group!` still
delegates to `criterion_group!`. Plain Criterion benches opt in with a
`config =` line and `MetricReport::iter` (documented in the gpui-bench
skill; wiring `rope` etc. is a follow-up). Reports are now emitted per
input for `inputs = ...` benchmarks, since per-iteration metrics blended
across inputs are meaningless. Criterion is bumped to 0.8.2, which also
randomizes stack alignment per sample.
**Verified** on Linux (hybrid Alder Lake) and an M-series Mac: every
mode on `inspector_render` and `editor_render`; counter denial via
seccomp on Linux; `bench_metrics` and `gpui` `bench_context` tests
including thread-scope inclusion/exclusion through real counters on both
platforms.
**Follow-ups:** counting-allocator metric (peak heap, allocations),
wiring non-GPUI benches, per-secondary baseline comparison.
Release Notes:
- [GPUI] Improved benchmarks to report retired instructions, cycles,
IPC, foreground-thread share, cache and branch misses, context switches,
and page faults per iteration alongside wall time, with
`BENCH_MEASUREMENT=instructions` or `=foreground-instructions` selecting
near-deterministic instruction counts as the Criterion-analyzed metric
for CI regression detection
e683fd7b46Anthony Eid committed on 9/27/2026, 5:37:29 PM· committed by GitHubparent4c841aa