gpui: Add randomized element tree benchmarks and fix Linux benchmark leaks (#64842)
Adds a seeded, randomized element tree for benchmarking GPUI rendering,
and a Criterion suite that measures what one frame costs under each
class of change. Running it on Linux exposed thread leaks in the
benchmark harness and in the Linux dispatcher that made every
`#[gpui::bench]` binary slow down as it ran; those are fixed here too.
## Randomized element tree (`gpui::randomized_element_tree`, behind
`bench-support`)
A tree of elements and entities generated from a `StdRng`, with wide,
deep, mixed and narrow topologies, entity-backed elements, and
interaction handlers. It supports the mutations a UI goes through (child
color and bounds, root bounds, insert, leaf removal, reorder, recoloring
a share of the tree either localized to one entity subtree or spread
out) and counts what re-rendered, so benchmarks can assert the renderer
did the work. `RandomizedElementTreeBounds` describes a family of trees
(topologies, element-count and density ranges) and `sample(&mut rng)`
draws one.
## Benchmarks (`gpui_platform/benches/randomized_element_tree.rs`)
Groups, each measuring frames in which exactly one thing happens: full
refresh (`Window::refresh`), unchanged (a notified root), leaf style,
leaf bounds, root layout, insert-remove, reorder, and a changing-share
grid (0/1/5/25/100% localized, 5/25% spread).
- **Families:** `any`, `wide`, `tall`, `mixed` and `dense-entities`,
each with at least 5% entities. Each family is a `#[gpui::bench]` input
seeded through the `StdRng` support from #64843, so ids read
`RandomizedTree/<group>/tree/<family>/seed-<n>` and `SEED`/`ITERATIONS`
choose trees as they choose `#[gpui::test]` seeds. The family name is
mixed into the RNG so overlapping families don't draw the same tree at
the same seed.
- **Reports:** each tree is described once with the heap its window
holds after the first frame. The bench report adds instructions,
allocations and frame timings from `bench_metrics`, and a final summary
prints each group and family's geometric mean of time and allocations
per iteration across seeds.
- **Run time:** two seeds per family (one for the changing-share grid)
and a 0.5 s warm-up with a 1 s measurement: 105 benchmarks in about 4.5
minutes on Linux. `ITERATIONS=6`, `--warm-up-time` and
`--measurement-time` give a longer run.
- **Checks:** every measured loop asserts that the tree actually
re-rendered, so a faster frame that skipped the changed element fails
rather than passing as an improvement.
```
cargo bench -p gpui_platform --features bench-support --bench randomized_element_tree [-- FILTER]
```
## Linux fixes for `#[gpui::bench]` and the dispatcher
On Linux the full suite reached thousands of threads and took about 40
minutes, each benchmark slower than the last. Measured frame times are
unchanged by these fixes.
- **A platform per routine call.** `#[gpui::bench]` builds a context for
every Criterion routine call, and each built a whole
`gpui_platform::current_platform(true)` just to take its text system.
`gpui_platform::bench_text_system()` now builds it once per thread.
- **A GPU device per window.** `WgpuHeadlessRenderer::new` created a new
wgpu device each time, about 65 ms per call. `WgpuContext::new_headless`
now shares one device per process, held in a static that is never
dropped (wgpu's queue reads its own thread-locals on drop, so a
thread-local cache aborted on thread exit), and creates a new one if it
is lost. Device errors are tracked per device, so a renderer joining a
shared device observes only errors raised from then on.
- **Dispatcher threads outlived their platform.** Dropping a Linux
platform left all 21 threads running. Workers block in
`PriorityQueueReceiver::recv`, which checked for dropped senders only
before waiting, and dropping the last sender never signalled the
condition variable. It now wakes every waiter (acquiring the queue lock
first so no receiver misses the wake-up), `recv` re-checks after each
wake, and the timer thread's event loop stops when its channel closes. A
regression test fails on the old code and passes with the fix. The
editor builds one platform per process, so it isn't affected in
practice; benchmarks, visual tests, and
`gpui_platform::background_executor()` are.
## Validation
On Linux:
- Full suite: 105 benchmarks in 262 s, none failing, with the thread
count flat (about 63 to 68) and live heap between routine calls flat (6
to 8 MB). Before the fixes: about 40 minutes, 2,315 threads 20 s into a
group.
- One frame per iteration: window draws equal iterations exactly, and
each iteration renders the root once and every element once.
- `gpui` (437, run on one thread, since 4 profiler and `bench_context`
tests share global state and also fail in parallel on `main`) and
`gpui_linux` (42) tests pass; clippy and fmt are clean.
On `main`, every class of change costs about the same as a full refresh,
and the changing-share grid is flat from 1% to 100%, because any notify
re-renders the whole tree. That's the curve retained rendering should
bend.
Release Notes:
- [GPUI] Added randomized element tree benchmarks, and fixed Linux
platform threads not exiting when the platform is dropped and
`#[gpui::bench]` creating a platform and GPU device for every iteration
---------
Co-authored-by: Mikayla Maki <mikayla.c.maki@gmail.com>
3135cdf8dcAnthony Eid committed on 9/28/2026, 4:17:17 PM· committed by GitHubparentc821538