Evidence boundary
The runnable tests and fixture outputs are synthetic host evidence. The authoring evidence report stores exact commands, compiler version, exit codes, fixture outputs, mutation results and SHA-256 source hashes outside the learner repository. Reproduce your own receipt with the commands in README and record the current source commit/hash, tool versions, actual results, and unresolved checks.
Label each observation as host-test, synthetic, firmware-compile, board-observed, or human-reported. An agent response is not a substitute for a command result. A cross-build does not establish upload success, sensor accuracy, electrical behavior, physical timing, serial enumeration, or indicator operation. No physical board observation has been performed for this project.
The four mutation exercises deliberately break isolated copies. A passing mutation harness means each bad variant was rejected while a fresh good copy still passed. Read the original failing assertion and the exact source diff before attributing a cause. Do not treat a compiler failure caused by a typo as proof of a behavioral test.
Behavioral discovery of AGENTS.md and the two skills must be recorded from real fresh agent sessions. The shipped prompts/rubrics describe the test; their presence is not a passing result. See skills. Windows guest builds and any future bench results require their own receipts rather than extrapolation from macOS.