Every claim that enters the Crucible now gets a public, timestamped machine forecast before the replication is designed or run. The forecast is committed to a public git history first, so it can be wrong in public — that's the point. Over time this becomes a prospective, contamination‑immune calibration record of machine skepticism about AI claims: an idea proposed in 2020 (arXiv:2005.04543) that, as far as we can tell, nobody has executed.
| Claim | P(reproduced) | Status |
|---|---|---|
| The compounding-error law. Chaining agents multiplies per-step error geometrically
(R = 0.95n → 77% at 5 steps, 60% at 10) — quoted as measured fact across
5+ vendor blogs, always without data. c001 · forecast 2026-07-10, committed before harness |
0.25 | replicating |
| The MCP tool-overload cliff at ~20 tools. Tool-selection accuracy "falls off a cliff"
past ~20 registered tools — a sharp universal threshold, per 4+ fresh posts (which quote
thresholds anywhere from 5 to 50). c002 · forecast 2026-07-10, committed before harness |
0.25 | queued |
| A skill-memory layer doubles agent win rate. AgenticSTS (arXiv:2607.02255) reports
3/10 → 6/10 wins in a long-horizon game — at n = 10 and p ≈ 0.37.
c003 · forecast 2026-07-10, committed before harness |
0.175 | queued |
P(reproduced) is the mean of two frozen model families (glm-5.2, deepseek-v4-flash; prompt v1, temperature 0), reading only the public claim card. All three currently bet against the trailing 0.80 reproduce rate — the forecaster thinks these particular claims won't survive measurement. Each forecast file, claim card, and the frozen prompt are in the public repo; the git push timestamp precedes any harness code for that claim.
CONTAMINATED in the repo and is never citable as skill. Only forecasts made
before their replication exists count.