proposal/14-toolchain.md

Toolchain: strata fmt, strata lint, strata-ls

The canonical formatter, the ledger-lint tier, and the language server. The section's actual thesis: three tools, one tree, one ledger, one diagnostic vocabulary — fmt owns the tree's fidelity, lint queries the ledger, and the language server renders both without ever growing its own approximation of anything. Principles in force: P-6 (determinism extends to tooling), P-9 (diagnostics are the product), P-7 (estimates never bear load), P-11 (agents and humans query one substrate). Charter answers from Veronica (2026-07-17) are marked SETTLED in the register.

strata fmt

1. Zero-config, day one (D-049)

No config file; no flags that affect output (only --check/--diff). Line width is 100 columns, baked (SETTLED — 80 forces ugly wraps on typed port declarations, 120 hurts side-by-side diff review, our dominant reading mode). The style versions with the edition (D-043), so style evolution is opt-in per component, never ambient — this replaces both rustfmt's option drift and Black's stable/preview churn. fmt-clean is a hard forge publish gate and never a build gate (SETTLED): strict at the library boundary, permissive during exploration — the Verilog landscape's lesson is that formatters without an ecosystem gate never take (Verible), and Strata negates the three structural reasons SV formatters failed: no preprocessor (D-048), a language born with its formatter, and a publish step to hang the gate on.

2. In-band steering (D-050, SETTLED)

The stance that survived large ecosystems is zero options plus steering read from the source itself (zig fmt lineage; Prettier froze its options and calls them a historical mistake; rustfmt's knobs produced per-project drift):

  • trailing comma → one-per-line; no trailing comma → single line if it fits;
  • blank lines are user-controlled grouping, preserved (collapsed to a max of one);
  • the anti-Black rule: strata fmt never inserts a steering token. Black's one deep regret (the magic trailing comma) came from the formatter writing trailing commas it then re-read as user directives — a one-way ratchet. We read steering; we never write it.

3. Layout doctrine: no vertical alignment (D-051)

stream {} / contract {} / port lists / match arms format one entry per line, single space around :/=, grouping via blank lines. Vertical alignment is the single largest source of diff churn and idempotence bugs in the gofmt/Verible record (one long field re-aligns its neighbors; Verible ships four alignment policies for one construct — option creep in miniature). Editors may render elastic alignment locally.

4. Formatted text is presentation; identity is structural (D-052)

Nothing in the compiler keys on formatted bytes — the Unison/Dhall move (hash normalized structure, not text), made mandatory by the documented history of formatter idempotence bugs (gofmt and zig fmt both, root cause almost always comment attachment). A formatter bug must never be a miscompile or a spurious version bump. The provenance-stability rule, normative:

For any source s accepted by the parser: elaboration of fmt(s) and s produce identical Architectural IR, identical ledger facts and identities, identical summaries, and identical memoization keys. Only recorded source spans may differ.

Mechanism: elaboration input is the AST, with a hard ban on comptime reflection over source text, columns, or comment contents (source layout is an undeclared input under P-6); ledger identity is path-based — (canonical path, generator identity, memo key, deterministic emission index) — with spans in a side table re-derived on every parse; summary hashes exclude trivia and spans. Corollaries: a reformat-only change is mechanically a patch with a reflexive equivalence certificate — no version bump (13 §3); P-9 diagnostics carry both the stable node ID (for agents — a repair survives a concurrent reformat) and the current span (for humans).

5. Verified invariants (D-053)

CI-fuzzed from day one: idempotence (fmt∘fmt = fmt), AST round-trip including comment count and anchors, cross-platform determinism, parse-error input returned untouched. Comments are trivia attached to anchor nodes (leading/trailing/dangling), never floating byte offsets. No fmt: off directive in v1 (a documented bug magnet; gofmt lives without one).

6. Scope (D-054)

strata fmt formats the whole package surface including forge.toml (canonical key order, one dependency per line) — one command formats a package, and registry-rendered manifests look identical to local ones.

strata lint

7. The soundness line (D-055)

R-LINT-1. A lint is never load-bearing. No judgment, obligation discharge, optimization-legality decision, summary compatibility check, or resolution outcome may depend on a lint's firing, silence, or suppression. Anything whose violation makes a design wrong is a judgment or an unsafe-category boundary, never a lint.

Corollaries: there is no correctness lint group — the group clippy had to deny-by-default is the group Strata deletes, because its contents are type judgments here; suppressing every lint changes no semantics, summary, or version, only the aggregate report; a lint discovered to be load-bearing is a bug filed against the type system, and its promotion to a judgment is a register event.

The triage of the commercial rule decks backs this up: SpyGlass/Verilator's high-value checks — unsynchronized CDC, width mismatch, latch inference, multiple drivers, undriven ports, combinational loops, reset-domain mixing, uninitialized state — all become judgments; a large second class (unsynthesizable constructs, full_case/parallel_case verification, X-propagation decks, sensitivity lists, blocking/nonblocking, ifdef hygiene, sim/synth-mismatch decks) dies entirely as artifacts of Verilog's weaknesses; what remains for lint is the genuinely advisory band — evidence gaps, physical estimates, missed optimizations, ecosystem hygiene. Full table in the research brief; condensed into the catalog below.

8. Architecture: declarative queries over the ledger (D-055)

A lint is a declarative query over ledger facts (07 §3) — the CodeQL model — never visitor code with assertion power: lints read facts, never write them, which makes third-party lints sandboxed by construction. Each lint declares the fact categories it reads; a pass that invalidates a category silences (never falsifies) downstream lints. The structural advantage over Infer/CodeQL-class tools: they reconstruct facts and spend their false-positive budget on analysis error; Strata's lints consume already-adjudicated facts, so the residual FP surface is judgment-of-taste only — exactly what levels and suppression-with-reason are for.

9. Groups, posture, surfacing (D-056, D-059)

Groups are a precision/maturity statement: suspicious (near-zero FP or it gets demoted), design, pedantic (opt-in), nursery (incubating — every new lint starts here), forge (manifest/ecosystem hygiene). Severity is consumer policy, never hard-coded by the lint; the default posture is warn, for every audience (SETTLED — one lint culture for humans and agents; release gates and agent policies may tighten explicitly, and their tightening is visible policy, not baked-in tiering).

Surfacing follows the Infer deployment lesson (diff-time review ≈70% fix rate; batch ≈ zero): default presentation is on the summary diff; full-tree results live in the aggregate report and gates. Every firing, suppression, and repair-acceptance is an experiment-database record; per-lint effective usefulness (repair-accepted ÷ fired, tracked per audience) drives Tricorder-style automatic demotion to nursery.

10. Suppression: expect with mandatory reason (D-057)

There is no bare allow. The only suppression form:

1expect(lint = cdc_reconvergence,
2 reason = "flags recombine ≥3 cycles after either edge; window proven in tb_corr",
3 evidence = Tested<corr_corpus>, // optional: reason backed by a fact
4 review_by = edition(2028)) // optional expiry, edition-keyed

Rules: reason is required (Rust proved optional reasons are never written; SpyGlass waiver files prove it at deck scale); a suppression is a ledger fact on the node it covers, pass-maintained like any fact — a pass transforming the node must preserve/transform/invalidate it, and an invalidated suppression re-arms the lint (the "alert when nearby code changes" recommendation, made sound); RFC-2383 semantics are native — an expect whose lint no longer fires is itself a diagnostic (suppression_stale), machine-applicably removable; scope is a node or component, never a file or tree; the suppression surface rolls up transitively in the aggregate report next to the unsafe surface, registry-queryable.

11. Repair applicability is verified, not declared (D-058)

Repair objects carry clippy's applicability taxonomy, but machine-applicability is checked: a repair is MachineApplicable iff applying the typed edit leaves the component summary bit-identical or strictly refining — the same summary diff that powers mechanical semver (13 §3). Clippy asserts applicability and has a recurring wrong-MA bug class; Strata verifies it with machinery that exists anyway. Contract-changing repairs (MI) state their impact in the structured-claim form agents already use (09 §2).

12. Publish gate: adjudicated, not clean (D-060, SETTLED)

forge publish requires zero unaddressed suspicious/design diagnostics — every survivor must be an expect-with-reason — and publishes the suppression surface (counts, reasons, ages) in the summary, mirroring unsafe-surface acknowledgment (13 §2). Hard lint-clean would convert every false positive into a publish blocker and breed the reason-less mass-suppression culture the FSE 2025 suppression study documents; no bar reproduces crates.io, where lint status is invisible at the registry.

13. Governance: three rings (D-060, SETTLED)

(1) The core team owns suspicious/design/pedantic; the only path in is nursery-first with telemetry-gated promotion. (2) Any package may ship scoped lints — sandboxed ledger queries applying to its own consumers (a protocol library legitimately knows what misuse of its contracts looks like). (3) Scoped→global promotion is a decisions-register event requiring cross-package telemetry. Third parties never add to suspicious: that group's near-zero-FP promise is the tier's credibility.

14. Starter catalog (v0, all entering via nursery)

Keyed on ledger facts; each ships a P-9 repair object. Condensed — MA = verified machine-applicable, MI = contract-changing candidate, PH = placeholder/plan:

suspicious: cdc_bridge_assumed_only (CrossClock guarantee at Assumed only → MI swap registry bridge / PH attach evidence) · cdc_reconvergence (two crossings sharing provenance, recombined without correlation contract → MI insert During window obligation) · arbiter_no_fairness_evidence (→ PH generated fairness property) · obligation_assumed_at_top (Tier-2 obligation re-exported as Assumed to build root → MI comptime_exhaustive where domain qualifies) · counterexample_unpromoted (DB counterexample with no regression test → MA emit test) · assumption_holder_distant (unsafe Ψ-fact discharged outside enclosing summary) · assumption_ownership_cycle (progress/availability assumptions form a closed owner cycle → MI add independent escape/rank / PH environment owner) · declared_relation_incomplete (memory-model contract omits required completeness without exporting an obligation) · certificate_stale (artifact dependency manifest invalidated by design/tool/target/mode/clock/calibration/lifecycle change → PH regenerate) · assurance_self_support (assurance graph appears in its own evidence closure; hard rejection, no suppression) · suppression_stale (→ MA delete) · suppression_expired (→ PH renew or fix).

design: likelihood_timing_secret_adjacent (likelihood-asymmetric timing on a path a labeled-secret value could ride — the D-071 side-channel caution as a lint; fires only where security labels exist, silent otherwise → MI equalize arm timing / PH declassification review) · fault_contract_unmutated (declared fault class has no injected mutant/campaign coverage → PH generate campaign) · health_detection_without_containment (detection bound exists but occurrence→containment window does not → PH add containment claim) · progress_bound_from_weak_fairness (finite N rests only on weak fairness → MI weaken to eventual / PH add quantitative premise) · observer_contract_unexercised (declared observer has no relational property or measurement plan) · estimated_fanout_exceeds (→ MI duplicate driver licensed by purity fact) · comb_depth_vs_clock (→ MI pipeline behind TransactionEquivalent) · buffer_capacity_unjustified (no Occupancy refinement → MI add named refinement / PH sizing experiment) · equivalence_overpinned (pinned CycleExact, consumers observe only TransactionEquivalent → MA relax, provably unobservable) · replicate_shareable (discharged exclusivity + low utilization → MI shared-serial variant) · dead_architecture (→ MA remove / MA expect for intentional taps) · reset_removable (no-read-before-first-write proven → MI reset-less register) · config_coverage_edge_gap (→ MA extend coverage corners) · chosen_param_unevidenced (D-035 choice without Measured backing → PH sweep experiment).

pedantic: latency_bound_slack (→ MA tighten to derived bound; refining ⇒ minor) · evidence_below_gate (→ PH verification-plan entry) · multicycle_unexploited (stability fact, no timing candidate → MI emit intended candidate plus mandatory post-synthesis binding plan, never an admitted exception) · assurance_residue_unowned (residue has no named disposition owner) · budget_headroom_low (→ MA raise declarative budget / MI restructure dominant generator).

nursery: stage_never_occupied (zero occupancy across all Tested corpora) · vendor_assumption_upgradable (registry indexes a refining proof artifact → MA import it).

forge: dependency_evidence_downgrade (lockfile update lowers depended-on evidence → MA pin) · suppression_surface_growth (transitive roll-up grew → PH review checklist by age × severity) · fact_dropped_at_boundary (emission's boundary ledger report shows a fact that neither baked, delegated, nor verified — 07 §2b → MI choose an exit / MA expect with reason).

strata-ls

15. The resident compiler; there is no second frontend (D-061, phase-1 gate SETTLED)

The field's verdict is unanimous: every ecosystem studied either paid for two frontends and paid again to merge them — rustc/rust-analyzer via mandated library-ification, Kotlin's K1 IDE reimplementation replaced wholesale by K2, Dart settling for shared-components, zls shipping no semantic diagnostics at all while Zig waits on its compiler-as-server endgame — or built one frontend from day one and never had the problem (Roslyn). Strata designs after the verdict: strata-ls is a long-lived daemon hosting the one compiler; strata build and strata-ls are two drivers over one library. Ships as a phase-1 gate, Tier A only (SETTLED — the one-frontend property is cheap on day one and brutal to retrofit; the language is used through the LSP from month one).

Interactive speed comes from machinery the proposal already has, not new query infrastructure: the summary firewall is the invalidation boundary; durability tiers map onto existing structure (registry summaries > workspace summaries > open bodies); body edits re-check one body against unchanged summaries; in-flight work cancels on edit. No salsa-class fine-grained memoization — Strata's declare-before-use modules, no preprocessor (D-048), and deterministic comptime (D-021) are exactly the language properties that make coarse map-reduce sufficient (matklad's "Against Query-Based Compilers", already load-bearing in 07 §6).

Normative invariants:

R-LS-1 (no second judgment path). Every diagnostic, completion, hover, or inlay hint is produced by the same checkers, over the same tree and summaries, as strata build. The LSP layer may select, schedule, and render; it may never approximate. Editor/batch divergence on identical input+config is a P0 bug class with its own CI oracle (fuzz both drivers, diff structured diagnostics). R-LS-2 (body-edit isolation). An edit whose exported summary is unchanged invalidates nothing outside that body; Tier-A latency is independent of workspace size. R-LS-3 (budgeted elaboration in-editor). D-044 budgets apply unmodified; exhaustion is a P-9 diagnostic naming the dominant call path, never a hung server. The editor may apply stricter interactive budgets, never looser (looser would accept what the build rejects, violating R-LS-1). R-LS-4 (config honesty). Every elaboration-derived artifact names its config point; bound-level facts are the only unlabeled facts. R-LS-5 (stable anchors). All diagnostics, repairs, and strata-ls/* responses reference nodes by path-based identity (D-052) plus current span — agent-held references survive reformats and concurrent edits.

16. One tree, three consumers (D-062)

The D-052/D-053 trivia-attached anchor-node AST is the single tree for parser, fmt, and LSP, with additive requirements: lossless for all inputs including invalid (text(tree) == input); error containment via ERROR-wrapped skipped tokens and zero-width MISSING tokens, with incomplete-but-present nodes for half-typed code (the resilient-LL discipline — the shape completion needs); incremental reparse with green-subtree reuse; and path-based identity defined over recovered trees too — ERROR nodes get deterministic paths, because broken code is the LSP's main diet and its diagnostics need anchors. Trivia stays anchored leading/trailing/dangling (the Roslyn/Swift model fmt already requires; rust-analyzer regrets not having it). Tree-sitter is explicitly not this tree (GLR recovery is unpredictable, upstream disclaims compiler-frontend use); its role is §18. The D-053 fuzz suite gains an LSP clause: every mutated/truncated input round-trips and yields diagnostics, never a parser failure.

17. Comptime in the editor: three tiers (D-063, D-064; active-config UI SETTLED)

The field's three strategies all have documented failure modes — zls's no-semantics gap (can't analyze without monomorphizing from main), clangd's heuristic dependent contexts, rust-analyzer's never-invalidate proc-macro cache hack — and each is pre-fixed by an existing Strata decision:

  • Tier A — every keystroke, no elaboration. All decidable judgments over the edited body against D-038 definition-checked bounds and imported summaries: width algebra (D-037 — normalization not evaluation is why width errors work without running comptime), phases, grades, protocol-catalog compatibility, with D-038's blame rule assigning bound-violation vs body errors. Completion, goto, hover from bound-level types. This is the tier zls structurally cannot have.
  • Tier B — active config, debounced, budget-bounded. One pinned instantiation kept incrementally elaborated, memo-warm via D-045 (a body edit re-runs only comptime calls whose memo keys changed — the caching rust-analyzer wants and can't have, sound here by D-021 construction). Yields concrete widths/latencies as inlay hints, ledger-fact hovers, elaboration diagnostics — all tagged with their config point (R-LS-4).
  • Tier C — on demand / background. Remaining coverage points ("check all corners" code lens), Tier-2 SMT, design lint — streamed via pull-diagnostics refresh, never blocking. OP-3's stable obligation outcome means "not yet known" is a first-class editor answer.

Active config (D-064, SETTLED as a visible picker): the "which instantiation am I looking at?" question every ecosystem answers with build-system folklore (clangd compile-commands, rust-analyzer cargo features) is answered here by the D-042 coverage set — declared, checked, evidence-carrying. Default is a designated-primary coverage entry (a new one-line publisher obligation); switching is a status-bar picker persisted in workspace state. Corollary that strengthens D-042's incentive story: the coverage set is no longer just a publish artifact — it is the IDE's menu, so authors gain editor features by declaring coverage.

18. Tree-sitter grammar and the Zed reference client (D-066; editor target SETTLED)

Zed is the reference client (SETTLED): its extension surface is small (extension.toml pinning the grammar rev, config.toml, .scm queries, one WASM shim wiring language_server_command), it forces the tree-sitter deliverable, and for this team the agent session (§20) is arguably the real first client; VS Code follows as packaging, agent-buildable against the documented extension surface.

One tree-sitter-strata repo under the project org (pre-empting Zig's four-competing-grammars fragmentation), per-editor query directories (highlights/brackets/outline/indents/injections/textobjects/runnables — Zed conventions; Neovim/Helix differ). The coexistence rule, normative: the compiler's parser is the sole source of truth, and CI runs the tree-sitter grammar over the compiler's full test/fuzz corpus (shared with D-053) — any compiler-accepted file producing ERROR/MISSING nodes, or compiler-rejected file parsing cleanly beyond declared tolerance, fails the build. Grammar updates ride the same PR as parser syntax changes. The grammar is presentation-only: no tool may derive semantic conclusions from it (corollary of R-LS-1). Strata's language properties — no preprocessor, no context-sensitive lexing (D-048) — are exactly what makes SV nearly un-grammar-able and Strata cleanly grammar-able.

19. Hardware-native editor features (D-065 phase tiers)

The features with no software analog are all renderings of ledger queries the batch compiler already answers — the LSP adds presentation, not analysis. Phase 1 (gated): daemon, incremental sync, Tier-A pull diagnostics, completion/goto/references/hover, semantic tokens with phase-modality modifiers (static/hardware/ghost/symbolic as token modifiers — ghost code renders dimmed, degrading gracefully in clients that ignore unknown modifiers), in-process strata fmt, code actions from P-9 repair objects (D-058-verified MachineApplicable → quickfix + isPreferred; contract-changing repairs behind the protocol's native needsConfirmation annotation — a delivery-side seatbelt on top of D-058's fix), the full P-9 object riding Diagnostic.data, tree-sitter grammar + Zed extension with corpus conformance in CI. Phase 2 (ungated, matures with the elaborator): Tiers B/C; ledger-fact hovers and inlay hints (inferred width, clock domain, latency/II, evidence tag — with label-part links to the defining contract); strata-ls/provenance ("why does this register exist" as an editor command), strata-ls/structureReport (P-2 cost lens), strata-ls/ledgerQuery; unsafe/suppression-surface lenses; coverage-corner code lens; strata-ls/traceAnchor mapping node IDs ↔ emitted signal paths for Surfer/VaporView-class waveform viewers and counterexample→source navigation (08 §5's loop, landed in the editor). Prior art stops almost exactly where this list starts: Verible LS/svls/veridian are syntax-level; VHDL-LS resolves names fast but elaborates nothing; only the commercial DVT IDE has elaboration-aware navigation.

Custom requests are namespaced strata-ls/*, each documented in an lsp-extensions.md with a CI doc-code hash check, and each new request is a batched register event (SETTLED — the extension doc is the contract agents build against; rust-analyzer's hash check exists because undocumented extensions rotted).

20. Agent access: same daemon, session-multiplexed (D-067)

strata-ls is multi-session by design (the Dart analysis-server shape; LSP's 1:1 assumption handled natively rather than by ra-multiplex-style bolt-ons): editors attach over stdio/socket, agents attach over a JSON session speaking standard LSP plus the strata-ls/* set — same checkers, same ledger, same memo-warm state. Agents get exactly what the wire already carries: P-9 objects with stable node IDs, verified applicability, evidence tags. A thin MCP adapter is phase-2 packaging. Session identity gives per-audience telemetry for free (§21). Telemetry consent (D-068): the event schema is designed now (it is the experiment DB's schema), external collection defaults off, first-party only through phase 2, revisited at registry launch.

21. The OP-5 measurement harness

OP-5 demands a worked multi-judgment diagnostic taxonomy "user-tested (or agent-tested: measure repair success rate given only the diagnostic)" — and until now had no delivery vehicle. strata-ls is it: multi-judgment failures render as one primary diagnostic naming its judgment, with relatedInformation carrying the cross-judgment facts (P-4's "no cross-judgment soup" as a rendering rule); every repair acceptance, rejection, or manual-fix-instead is an experiment-database event keyed by (judgment/lint, repair, audience) — the same repair-accepted÷fired stream D-059 uses for lint demotion, extended to judgment diagnostics. The Infer deployment lesson (diff-time ≈70% fix rate vs batch ≈0) predicts the editor channel dominates repair uptake, making it the right place to instrument first. OP-5's "resolution must deliver" gains: per-judgment repair-success dashboards from strata-ls telemetry by end of phase 2, cited by the P-9 ship-gate for new judgments.

North star

  • Scoped-lint authoring as a first-class forge artifact kind, with the ring-2/ring-3 promotion pipeline instrumented end to end.
  • Formatter-aware merge (structural three-way merge over the canonical AST — the D-052 identity layer makes this tractable).
  • Lint-usefulness telemetry published per lint on the registry, closing the Tricorder loop ecosystem-wide.
  • VS Code extension (packaging over the stable lsp-extensions.md; agent-buildable).
  • Waveform-viewer co-development: strata-ls/traceAnchor adopted by an open viewer so click-signal→go-to-driver ships outside commercial IDEs for the first time.
  • Editor-embedded schedule/pipeline visualization rendered from schedule witnesses (03 §4).

Open problems touching this doc

Consumes OP-6 (suppressions are pass-maintained ledger facts — ledger erosion would silently disarm re-arming) and feeds OP-5 — strata-ls is OP-5's designated delivery vehicle and measurement harness (§21): repair objects are the diagnostic vocabulary's highest-volume consumer, and per-judgment repair-success telemetry is how the P-9 ship-gate gets its data.