Method

How a number gets onto this site

Most of what I run is small, quick and done by me alone. That is fine for finding out what is worth testing properly. It is not fine to present as settled. These rules keep the two apart.

Evidence classes #

Every record carries one or more of four marks. The mark says what kind of thing the number is before you read the number.

Measured

Something I ran and recorded: a model scored on audio, a wall meter read during a run. It describes that run. It does not describe other models, other days or other machines.

Modelled

A formula or a calibrated estimate. The inputs may be measured, but the output is a calculation. When a model and a meter disagree, the meter wins and I say so.

Simulated

Data a script generated from settings I chose. Useful for showing how a method behaves under known conditions. It cannot show how real systems behave, and a simulation can never confirm its own settings.

Proposed

A question and a plan, with nothing run. I list these so you can tell a result from an intention.

What every record shows #

  • The date and time of the run, in UTC.
  • The sample: how many models, items, seeds or runs.
  • The method in one line, and the results file the numbers came from.
  • A short statement of what the result does not show.
  • A link to the full paper, where one exists.

If I cannot fill in one of these, the number does not go on the page. I do not publish a figure without a date, a sample and a method. Where a result has no confidence interval, the page says so and treats the figure as a point estimate.

Where numbers come from #

The pages do not hold typed measurements. A script reads each figure from the results file and records the SHA-256 of that file. If a results file changes, the site fails to build until the numbers are read again. That is how I stop a page and its data from drifting apart. It does not make the data correct. It makes the page match the data.

Each figure names the results file it was read from, and those files are published with their digests. The audio dataset and its manifest are public on Hugging Face, so the file checks can be repeated by anyone.

Review status #

Nothing here is peer reviewed, and I know of no outside reproduction of any result. I run a checklist script before I publish. It catches formatting and consistency mistakes. It is not a review, and it did not catch most of the errors on the corrections page.

I will not use words such as proven, verified, cleared or validated for my own results. They imply a checker, and there is none.

Naming #

Third-party models appear by size and kind, for example “a 39M-parameter English-only model”, not by product name. Hardware is named where the specification matters. I am the only author, and Echo is my own project.