Agent memory · measurement

An agent's past lessons helped when outside outcomes wrote them, not when a ranker picked them

September 28, 20263 min readAgent memory · Memory poisoning · Outcome credit · Taste-Bench · inspeximus
The takeaway

On 160 Taste-Bench engineering forks, random same-repository lessons lifted Claude Opus 5.5 from 0.594 to 0.750; similarity ranking added 0.019 more, inside noise. Lessons written from outside outcomes helped, self-judged ones did not. Probe and per-fork results included.

Lessons that carried a confirmed outcome lifted Claude Opus 5.5 from 0.594 to 0.750 on engineering judgment forks, even when we picked them at random. Picking them by embedding similarity raised that to 0.769. That extra 0.019 has a 95% interval of -0.031 to +0.069, so this run cannot tell it from zero.

We set out to show that recalled decision memory makes an agent choose better. The result is narrower: what helped was lessons written from outcomes confirmed outside the agent. How we ranked them mattered little on these forks.

What we measured, in order

We used the engineering split of Taste-Bench (Pan et al., arXiv 2609.25804): 390 forks from real coding trajectories. Each fork has two candidate next steps, and the label is the one whose branch had the better recorded outcome. The model sees both options in both orders, and a fork counts only when both answers are right, which controls for position bias.

A lesson is one sentence from a fork in another task: "choosing X over Y led to the better outcome". We never recall a lesson from the task under test.

Lessons in the promptBoth orders correct
None0.594
Random, same repository, same surface cues0.750
Top 3 by embedding similarity (inspeximus recall)0.769
No model: a two-line rule (prefer the isolated or longer option)0.669

Random same-repository lessons add 0.156 (+0.087 to +0.225). Similarity ranking adds 0.019 on top. At 160 forks, the run can rule out a ranking gain above about 0.07, not a smaller one.

Where the outcome came from mattered

We let Qwen write its own lessons from its own earlier choices, in two ways:

This is consistent with Huang et al. on reasoning (arXiv 2310.01798): without outside feedback, models struggle to correct themselves, and their performance sometimes degrades. ReasoningBank (arXiv 2509.25140) is a counter-case: self-judged memory helped on web tasks, where success is easier to check than on a judgment fork.

Where memory poisoning fits

Memory that learns from outcomes is exposed to memory poisoning: a false outcome, written into the store. We gave every lesson a flipped twin, a deliberately extreme store where about half of every recalled set is wrong. This was a separate Opus run at default effort, where the clean memory arm scored 0.825. Plain retrieval fell to 0.637. Recalling only lessons whose outcome was credited from outside restored 0.825.

That last number shows what the gate does with the signal, not that it detects an attack: in this harness the outside credit is the dataset's own label. An earlier post showed the same boundary in stylized demonstrations: a defense that credits self-graded outcomes falls to an attacker who grades their own poison as a success.

What would have proved us wrong

Limits

What this means for inspeximus

The part of agent memory that carried the result is the outcome, and who confirmed it. inspeximus supports that, and it is opt-in. In the Python API, recall(influence_only=True) returns only corroborated memories. With credit_requires_warrant=True, credit counts only when it carries a warrant. remember(project=...) and recall(project=...) scope memory to a project, and unstamped records stay visible. On this benchmark, better similarity ranking was not where the value was.

Every number here is printed by research/probes/taste_outcome_not_ranking.py from the per-fork results beside it.

← More writing from Agora