Frozen evidence
Every release names the dataset, split, licence, evaluator version, prompt or rule, and execution date.
Measured reports
Reproducible tests for the gates between a promising model and a production system. Raw results and methods included.
Measured releases
Frozen datasets, explicit selection rules, tail latency, limitations and raw results.
REPORT 01 · EXECUTED 2026-08-23
2,675 held-out responses. Precision, recall, F1, false-negative rate, local latency, raw decisions and the configurations that failed.
REPORT 02 · EXECUTED 2026-08-24
5,000 COCO validation images. Twelve configurations across M4, Core ML, L4 CUDA and C4 ONNX, with 2,400 timed inferences and capacity scenarios.
REPORT 03 · EXECUTED 2026-08-24
One million successful requests through a public GCP edge, 13.45 ms p95, zero HTTP errors, forced database traffic and an explicit production cost boundary.
REPORT 04 · EXECUTED 2026-08-28
Six injected failure modes, three models, 219 runs. Nothing was fabricated. Two of the three models never retried a tool that explicitly asked to be retried.
Publication standard
Every release names the dataset, split, licence, evaluator version, prompt or rule, and execution date.
We publish false negatives, failed configurations and limitations—not only the metric that makes a gate look useful.
HTML is ungated. Raw results, methodology and executable reproduction files are downloadable without an email.
Release queue
04AIoT anomaly and offline-recovery benchmark — planned with NASA C-MAPSS
05State of AI Project Readiness — only after 100 estimator submissions and 20 responses per segment