Results files

The numbers behind every record

Each figure on this site is read from one of these files. They are the raw results of the runs, one JSON file each, with a SHA-256 you can check.

Licence CC BY 4.0Self-published, not peer reviewed

The files #

Sizes and digests are of the files as published. index.json lists the same data for scripts.
FileWhat it holdsRun (UTC)SizeSHA-256
e1-hallucination-ledger.json
Record
E1: words written for silence, noise and squelch, with and without an energy gate2026-10-10 17:48:31Z109 KB759cf097edea2ec04becfebbbe2a024d1b356069fe961f19eeaa73c4376252d8
e2-telephone-grid-ledger.json
Record
E2: six speech models through six audio paths and five noise levels, every transcript2026-10-10 18:34:02Z881 KB49390df45c82d52fa4278a61a935dee2d0b3a001a614fd3092205a6443943a16
e3-mast-coding-ledger.json
Record
E3: 48 development attempts coded for failure modes, single agent against two coordination setups2026-10-10 17:40:21Z28 KB4f7b2ae0f69237e36f4422bc2453a59d95a83cec225bb3b8f4b210f08b627f7c
sovereign-apu-inference-ledger.json
Record
Energy: wall-meter runs and calibrated model rows on one unified-memory machine2026-10-10 23:26:56Z29 KBfa05c73f72adeb028ce7d0d0108df21f7b1c9da456d0a80d62ba92f74fe7eade
vec-scr-evaluation-ledger.json
Record
Verifier simulation: error covariance and budget sweeps on a seeded synthetic pool2026-10-10 21:51:06Z27 KB180120350f19b243d6ab3da83498ee1022c70eeae744703f87458896db940d8c

Check a figure #

The first record says mean word error was 59% with no added noise and 129% at 0 dB. This reads the same figures from the published E2 file. It needs only Python.

# mean word error across all models, by noise level
curl -sO https://roamingpigs.org/results/e2-telephone-grid-ledger.json
python3 - <<'PY'
import json, statistics
d = json.load(open("e2-telephone-grid-ledger.json"))
for k in ["clean", "snr_20", "snr_10", "snr_0", "snr_-6"]:
    print(k, round(statistics.fmean(m["by_snr"][k]["wer"] for m in d["models"]) * 100, 1))
PY

Expected output: clean 58.9, snr_20 60.1, snr_10 80.4, snr_0 129.4, snr_-6 157.7.

Energy per token is the same kind of check: in the energy file, divide wall_power_mean_watts by tokens_per_sec for each entry in live_multi_model_sweep.

What was removed #

Product names of models are replaced by size and kind, for example “39M English-only”. Local file paths, machine addresses and host names are removed. The export checks that no number changed: the sorted list of every number in each file is identical before and after. It also fails if a product name or a path is left in.

What is not here yet #

  • The scripts that produced the results. They are not published. Until they are, you can check the arithmetic on these files but not re-run the experiments.
  • Model identities. Because names are replaced by size and kind, you cannot yet re-run the exact models. This limits independent reproduction, and I know it.
  • The audio. It is on the dataset page. Note that the speech-model run scored a different set of utterances from the released files.