Snowbound writes into other people's notebooks, often while their copy of OneNote is writing to the same file. A bug doesn't just crash an app. It can corrupt a shared notebook, or silently drop a paragraph someone else typed. The testing strategy follows from that. Never grade the code with itself. Every important claim is checked against something the code under test didn't produce: a real OneNote, a separate model, a second path through the system, or a disk that loses bytes on purpose.
This essay explains where the confidence comes from. The how-to for running each lane is in tools/TESTING.md and the crate READMEs.
The only authority on "OneNote accepts this" is OneNote. The lab (tools/w7)
runs OneNote 2010 in disposable QEMU clones of a sealed Windows 7 image. An
agent inside each clone runs AutoHotkey and PowerShell (OneNote's COM API),
returns screenshots and files, and the clone is discarded afterwards.
Every storage feature goes through the same loop:
| 1 | observe author the edit in OneNote (COM, or driving its UI) and dump what it stored |
| 2 | write make Snowbound store the same thing through ops; export a candidate notebook |
| 3 | cold-open a fresh clone with a fresh OneNote cache opens the candidate |
| 4 | ─► integrity check passes, XML export, screenshots, PDF where layout matters |
| 5 | gate a VM-free test compares the capture with what the candidate meant to say |
| 6 | install candidate + capture become a corpus row (corpus/<feature>/…) |
Cold matters. A warm OneNote has its own cached copy and can hide a file it would reject. Some failures appear only at scale. OneNote once silently dropped elements in about one build in five because of an identity choice its integrity check accepted, and it refuses revision chains past a certain depth. Gates therefore include large mixed candidates, not only one small row per feature.
The corpus is what makes this sustainable. Each row keeps the native capture
next to the candidate, and its tools/test_*.py gate checks it without a VM,
on every run, forever. Re-running the VM step is needed only when the bytes
Snowbound writes change.
Most checking needs no VM. It works by making independent views of one edit agree:
onestore::op::model interprets ops on the page model
without going near bytes. After a random edit is applied to a Section,
sealed and read back, the page must equal what the model predicted.
Refused edits must leave the page unchanged.crates/canvas/tests/ops_differential.rs).Durability claims are tested by making things fail at every point that matters.
tools/smb-proxy.py sits between the client and a lab
Samba server and withholds a chosen request or response. The matrix cuts
every message of a commit. Each outcome must keep its promise:
NotCommitted really didn't publish, Committed really did, and Unknown
cases really went either way, with the confirmation that follows finding
out which.The collaboration lab puts several OneNote 2010 clients and several Snowbound writers and readers on one section in a Samba share. They make random edits through outages, reconnects, offline stretches and OneNote's own maintenance (Optimize rewrites the file underneath everyone). The pass criteria are strict:
fuzz/ is its own cargo-fuzz workspace. It covers parsing (storage, property
streams, revisions, documents), the commit protocol with interruptions, edits
of many kinds, the page model, protected sections, the offline queue, and the
canvas editor's state machine. The parsers see arbitrary bytes, and the
writers see arbitrary sequences of edits whose results must still validate.
Tests fall into two kinds, and the suite is being split to match:
cargo test
over the workspace, clippy, and the Python gates over retained captures.
tools/check_public.py runs all of it from a clean checkout, with no private
notebooks, VMs or credentials.CANVAS_SWEEP_* and OPS_SWEEP_SECTIONS) and
are meant to run for days or weeks on dedicated machines. What they find gets
reduced to a small deterministic test in the first tier.Lab lanes (real OneNote, Samba VMs, the proxy) sit beside both tiers. In Rust
they appear as ignored tests that name the environment they need, and the rest
live in Python harnesses under tools/.
The project is strict about what a result establishes, and that strictness is itself a source of confidence: