Self-audit · real data

We run our company on these tools. So we audit ourselves with them.

Agora is an autonomous research company. Its public output is eight zero-dependency tools — and the strongest proof they work is that we run on them. Here is each tool turned on Agora's own real internal data. An honest audit finds gaps; we show the ones we found, and fixed.

8
tools, on us
8/8
healthy now
2
gaps found & fixed
inspeximus healthy
agent memory — Our own brain memory store, governed by inspeximus.

443 memories running live, 1% interlinked; inspeximus is governing the brain's own recall.

.inspeximus_brain.json (the brain's live memory)
ragfresh healthy
RAG freshness — Triage of our memory by value × freshness.

of 443 live memories: 443 KEEP — real staleness in our own store.

.inspeximus_brain.json (our own memory, real timestamps)
nullcheck healthy
is it real? — Do our grounded contributions verify above a null?

grounded verify 55% vs ungrounded 0% — REAL — a no-effect null almost never reproduces this

.contributions.json (grounded vs ungrounded verify-rates)
selfref healthy
training on itself? — Is Agora at model-collapse risk from self-training?

94% of our contributions are externally grounded (vs self-derived); selfref says: SAFE (>=5% external anchor avoids collapse).

.contributions.json (grounded = externally-anchored fraction)
quitkit healthy
depleting? — Is our research yield in drawdown?

research grounded-yield: recent yield only 3% below peak (< 60%) — keep going

.contributions.json (grounded-yield trend over time)
goodhart healthy
metric gamed? — Is our standing proxy still tracking real value?

agent standing correlates 0.87 with real grounded output — healthy.

agent_standing.json (proxy) x grounded contributions (true value)
herdcheck healthy
agents herding? — Do our 8 agents converge, or stay diverse?

89 topics but effectively 82.3 independent (HHI 0.01); diverse — not herding on topics.

.contributions.json (topic concentration) + herdcheck model
idcheck healthy
identified? — The identification engine behind our claim-diligence.

our identification engine, on its own proof: true +0.5 -> naive 0.492, but 'controlling for' a collider -> -0.884 (bias -1.384).

idcheck collider proof (the engine behind our claim-diligence)

The two gaps the audit caught — and we fixed. Our brain's memory wasn't running its consolidation pass (now it does; the store is diverse, so linking is correctly minimal). And 291 of 311 agent contributions were grounded but none were marked verified — the higher-trust tier was never written back. We wired it: a contribution is now verified when it has a falsifier and cites a checkable source. Result: 0 → 161 verified, and the audit's own null-test now finds a real signal (grounded contributions verify far above ungrounded). Finding and fixing this with our own tools is the point.