Crucible Live

Forecasts before replications.

Every claim that enters the Crucible now gets a public, timestamped machine forecast before the replication is designed or run. The forecast is committed to a public git history first, so it can be wrong in public — that's the point. Over time this becomes a prospective, contamination‑immune calibration record of machine skepticism about AI claims: an idea proposed in 2020 (arXiv:2005.04543) that, as far as we can tell, nobody has executed.

3
Open forecasts
0
Resolved
0.80
Trailing base rate

Live forecasts

ClaimP(reproduced)Status
The compounding-error law. Chaining agents multiplies per-step error geometrically (R = 0.95n → 77% at 5 steps, 60% at 10) — quoted as measured fact across 5+ vendor blogs, always without data.
c001 · forecast 2026-07-10, committed before harness
0.25 replicating
The MCP tool-overload cliff at ~20 tools. Tool-selection accuracy "falls off a cliff" past ~20 registered tools — a sharp universal threshold, per 4+ fresh posts (which quote thresholds anywhere from 5 to 50).
c002 · forecast 2026-07-10, committed before harness
0.25 queued
A skill-memory layer doubles agent win rate. AgenticSTS (arXiv:2607.02255) reports 3/10 → 6/10 wins in a long-horizon game — at n = 10 and p ≈ 0.37.
c003 · forecast 2026-07-10, committed before harness
0.175 queued

P(reproduced) is the mean of two frozen model families (glm-5.2, deepseek-v4-flash; prompt v1, temperature 0), reading only the public claim card. All three currently bet against the trailing 0.80 reproduce rate — the forecaster thinks these particular claims won't survive measurement. Each forecast file, claim card, and the frozen prompt are in the public repo; the git push timestamp precedes any harness code for that claim.

The protocol

  1. Intake. A claim card (claim + source only, no replication design) is written. Selection criteria: fresh and getting attention, computable as a minimal single-machine model, and FAILED must be a live possibility. Rejected claims are logged too.
  2. Forecast. The frozen forecaster reads only the claim card and outputs P(reproduced) with a reason. The trailing base rate is frozen into the file as the comparator.
  3. Commit. Card + forecast are pushed to the public repo before any harness exists. The push is the timestamp.
  4. Replicate. The standard Crucible severe-test pipeline runs — smallest faithful computational model, runnable receipt. Forecasts are append-only; nobody edits them after seeing data.
  5. Resolve. The verdict lands in the ledger; the forecast gets its Brier score in a new commit.

What we will and won't claim

No skill claims before 60 resolved prospective forecasts. A retrospective pilot on our 33 already-published verdicts beat the base rate handily (Brier 0.095–0.109 vs 0.231) — but those verdicts have been public since May 2026 and may sit in training data, so that pilot is labeled CONTAMINATED in the repo and is never citable as skill. Only forecasts made before their replication exists count.
Known limitation, declared upfront. The same organization selects the claims, makes the forecasts, and runs the replications. The intake rules, the rejected-claims log, and append-only forecasts mitigate this; base-rate endogeneity (we deliberately hunt claims where FAILED is live) cannot be eliminated — which is why the comparator is the trailing base rate, not 50/50. Third-party claim submissions remove the selection loop entirely: submit a claim.