It generally applied the stored memory correctly, using it to avoid repeating basics and to produce useful handoffs, but one reply was slightly more structured and polished than the user’s strict short, direct preference.
What was measured
Correct Application
Checks whether the agent actually uses the retrieved memory correctly.
decisive for this rankingtransformation
For agent memory, it is not enough to retrieve facts; the system must help the agent use them correctly in task execution. (2 of 3 judges)
What was given, what came back
Input — what we sent
No input — this is a capability finding
The observation is about the tool itself rather than one test input, so there is nothing to show on this side by design.
Output — unretouched
No output artifact
The verdict rests on the tester's written observation alone — no file was captured for this cell.
Provenance
- Observation
- da199457-750c-4352-9321-8ef8aad3c587
- Evidence run
- 6e31afbb-34d7-459a-b688-68ef76fc615a
- Study
- Memory for AI Agents
- Research task
- 86ba16xrp
- Tested at
- not recorded
- Source
- aggregate-synthesis
- Evidence state
- observed
- Proof shown
- no artifact
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "hindsight"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Correct Application
No other tool was measured on this criterion for this input.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com