The autonomous engineering layer above the deterministic substrate. Principles in force: P-8 (models outside the trusted base), P-9 (diagnostics are the product), P-11 (agents build Strata with Strata).
The documented failure modes of LLM RTL generation map one-to-one onto what Strata's substrate eliminates or catches (D-028 research):
Trusted: parser, typechecker, comptime evaluator, IR verifier, lowering passes, contract compiler, reference-model interface. Untrusted: every model output — code, properties, rewrites, strategies, analyses. Admission paths: typechecking; translation validation / equivalence checking for rewrites (07 §4); vacuity + mutation scoring for properties (08 §6); simulation/formal/FPGA evidence for behavioral claims. An agent claim without evidence tags is not reportable (P-7).
Agents read and write: typed AST, architectural IR, contract IR, obligations, schedule witnesses, experiment definitions — not raw structural netlists. A proposed change is a structured claim:
| 1 | architectural changes: none |
| 2 | schedule changes: +1 stage between priority levels 2–3 |
| 3 | contract impact: latency +1; throughput unchanged; |
| 4 | CycleExact invalidated; TransactionEquivalent preserved |
| 5 | required validation: typecheck, schedule check, protocol refinement, local LEC |
The compiler rejects inconsistent claims before expensive validation runs. Diagnostics close the loop: every rejection is a structured object (violated judgment, facts, provenance, candidate repairs — P-9), turning repair into targeted search over the repair candidates rather than regeneration.
As in rough-1.md §26, kept: research, specification, architecture, implementation, verification, formal, synthesis, debugging, and review agents — with the review agent holding rejection authority and the whole set sharing one workspace protocol: claims in, evidence-tagged results out, everything logged to the experiment database. Role boundaries follow the summary system (07 §6): an agent's blast radius is the components whose summaries its change invalidates, which is what makes parallel agent work safe.
Every attempt — successful or failed — is a structured record: design hash, diff, config, rationale, compiler results, properties checked, formal results, coverage delta, corpus, synthesis/FPGA measurements, fault model and injection sites, observer/model choice, assumption owners, artifact/tool/target pins, expiry events, failure traces, review verdict, retention decision, and uncovered residue. Serves three functions: agents don't repeat rejected work; discovered invariants and counterexamples are reusable assets; the FPGA loop's measurements land as Measured evidence attached to designs, not as prose in a report.
The verification agent inspects semantic coverage → explains generator limitations (in terms of the typed strategies, which it can read) → proposes new rules/strategies → validates them (do they reach the gap?) → measures incremental value → retains only what pays. Property proposals go through the §1 admission path. This loop is the platform's compounding asset: the corpus and property set improve monotonically under mutation scoring.
Agents may propose assumptions, fault models, declassification policies, physical exceptions, or evidence links, but cannot self-approve them. Every assumption names an owning component or environment boundary; cycles in which A and B each assume the other's progress are rejected. Every artifact records invalidators—source/summary change, tool or target change, clock/mode/netlist change, calibration expiry, mapping/power/reconfiguration epoch death. The review agent sees a release diff over guarantees, assumptions, evidence axes, expiry, and residue, not a synthetic pass/fail score.
An assurance agent may assemble and explain the custody graph and identify missing links. It cannot cite that graph as proof, upgrade Tested to Proved, or hide incompatible evidence axes. Independent deterministic checkers, where available, supply diversity evidence; model consensus does not.
Consumes OP-5 directly — the measurable success criterion proposed there (agent repair success rate given only the structured diagnostic) should be a tracked platform metric from phase 2 on.