Results files
The numbers behind every record
Each figure on this site is read from one of these files. They are the raw results of the runs, one JSON file each, with a SHA-256 you can check.
Licence CC BY 4.0Self-published, not peer reviewed
The files #
| File | What it holds | Run (UTC) | Size | SHA-256 |
|---|---|---|---|---|
| e1-hallucination-ledger.json Record | E1: words written for silence, noise and squelch, with and without an energy gate | 2026-10-10 17:48:31Z | 109 KB | 759cf097edea2ec04becfebbbe2a024d1b356069fe961f19eeaa73c4376252d8 |
| e2-telephone-grid-ledger.json Record | E2: six speech models through six audio paths and five noise levels, every transcript | 2026-10-10 18:34:02Z | 881 KB | 49390df45c82d52fa4278a61a935dee2d0b3a001a614fd3092205a6443943a16 |
| e3-mast-coding-ledger.json Record | E3: 48 development attempts coded for failure modes, single agent against two coordination setups | 2026-10-10 17:40:21Z | 28 KB | 4f7b2ae0f69237e36f4422bc2453a59d95a83cec225bb3b8f4b210f08b627f7c |
| sovereign-apu-inference-ledger.json Record | Energy: wall-meter runs and calibrated model rows on one unified-memory machine | 2026-10-10 23:26:56Z | 29 KB | fa05c73f72adeb028ce7d0d0108df21f7b1c9da456d0a80d62ba92f74fe7eade |
| vec-scr-evaluation-ledger.json Record | Verifier simulation: error covariance and budget sweeps on a seeded synthetic pool | 2026-10-10 21:51:06Z | 27 KB | 180120350f19b243d6ab3da83498ee1022c70eeae744703f87458896db940d8c |
Check a figure #
The first record says mean word error was 59% with no added noise and 129% at 0 dB. This reads the same figures from the published E2 file. It needs only Python.
# mean word error across all models, by noise level
curl -sO https://roamingpigs.org/results/e2-telephone-grid-ledger.json
python3 - <<'PY'
import json, statistics
d = json.load(open("e2-telephone-grid-ledger.json"))
for k in ["clean", "snr_20", "snr_10", "snr_0", "snr_-6"]:
print(k, round(statistics.fmean(m["by_snr"][k]["wer"] for m in d["models"]) * 100, 1))
PYExpected output: clean 58.9, snr_20 60.1, snr_10 80.4, snr_0 129.4, snr_-6 157.7.
Energy per token is the same kind of check: in the energy file, divide wall_power_mean_watts by tokens_per_sec for each entry in live_multi_model_sweep.
What was removed #
Product names of models are replaced by size and kind, for example “39M English-only”. Local file paths, machine addresses and host names are removed. The export checks that no number changed: the sorted list of every number in each file is identical before and after. It also fails if a product name or a path is left in.
What is not here yet #
- The scripts that produced the results. They are not published. Until they are, you can check the arithmetic on these files but not re-run the experiments.
- Model identities. Because names are replaced by size and kind, you cannot yet re-run the exact models. This limits independent reproduction, and I know it.
- The audio. It is on the dataset page. Note that the speech-model run scored a different set of utterances from the released files.