Four more kinds of model can be measured, the report is rebuilt, and every run
now records what makes it comparable to somebody else's box.
Every run carries its provenance
A run taken without this is a run that cannot be compared tomorrow. Each file
now has a provenance block with a schema number on it.
Per run: the host that drove the benchmark (os, release, version, arch,
python, logical CPUs, physical RAM, machine model), the transport (plane,
gateway port, vhost if used, and the round trip measured before any work
starts), the device (name, model, serial, TiinyOS, service version, NPU
unit budget), and the tool (bench version, suite version, envelope schema,
git commit). A benchmark driven from a Mac over a cable and one driven from a
Windows box over Wi-Fi are not the same measurement, and until now nothing in
the file said which.
Per model rather than per run, because a sweep loads and unloads as it goes
and the conditions the fifth model met are not the ones the first met: the
resident set with each model's NPU units at the moment that model's tests
begin, units used and free, whether another process held the device lock, the
sampling parameters actually sent, which tests ran and which were skipped, and
the served context window. The lock is read and never taken, because a
benchmark that fought for it would change the thing it measures.
Temperature and power are recorded as null with a reason rather than omitted. …