Because the prior cache-clear memory was not used, the drafted reply recommended 'Clear Cache and Cookies' again, which directly repeats one of ACME's already-tried troubleshooting actions and violates the client's request to avoid repeated troubleshooting.
What was measured
Correct Application
Checks whether the agent actually uses the retrieved memory correctly.
decisive for this rankingtransformation
For agent memory, it is not enough to retrieve facts; the system must help the agent use them correctly in task execution. (2 of 3 judges)
What was given, what came back
Test input: Client Relationship Memory · text · group: memory-for-ai-agents
Input — what we sent
The exact prompt
Session 1: Use this under account_id: client_acme_001 ACME is a client using our AI support assistant. They prefer clear next steps and do not like repeated troubleshooting. Their team already tried password reset, clearing browser cache, and switching browsers. The issue is still happening only for users with SSO enabled. Session 2: Use this under account_id: client_acme_001 ACME came back today and said: "Our users still cannot log in with SSO. What should we try next?" Draft a support reply that respects what they already tried and moves to the next useful step. Session 3: Use this under account_id: client_acme_001 Update the client memory: ACME is no longer using SSO for this rollout. They moved to email-password login for the first launch. Do not keep treating SSO as the active issue unless they mention it again. Session 4: Use this under account_id: client_acme_001 ACME says: "Some users are still unable to log in during launch testing." Draft the next support reply. Session 5: Use this under account_id: client_beta_002 BetaCorp is a new client. They say: "Our users cannot log in for the first time." Draft the first support reply for BetaCorp.
A client-support memory test that checks whether the assistant remembers prior troubleshooting, handles a changed rollout context, keeps client scope isolated, and avoids leaking one client’s history into another client’s support reply.
Why this input is hard
- · client history recall
- · avoid repeating troubleshooting
- · update handling
- · scope isolation between clients
- · support response continuity
Output — unretouched

Also checked on this input — same tool, 4 other criteria
Memory Capture Quality✓ WorkedThe tool captured the client boundary and the durable support context in one memory set: ACME's identity, the already-tried steps (password reset, browser cache clear, browser switch), the preference against repeated troubleshooting, and the fact that the issue was SSO-specific.Relevant Retrieval✗ FailedIn the next support turn, the reply panel showed 'No memories retrieved yet,' so the assistant did not surface ACME's prior troubleshooting history and instead repeated a cache-clearing step that had already been tried.Scope Control✓ WorkedThe tool kept client memories isolated: the separate BetaCorp container showed 'No memories retrieved yet' and did not inherit ACME's SSO history, troubleshooting steps, or support preferences.Update and Correction Handling⚠ StruggledAfter ACME's rollout context was updated away from SSO, the UI still showed a '0 memory used' tag while the side panel continued to surface the old SSO memory at 71% and then 79% relevance on later replies, so it is not possible to confirm from the UI that the update reliably superseded the stale context.
Provenance
- Observation
- 25402b97-0e10-4fd2-bf59-da04f19e9704
- Evidence run
- 6e31afbb-34d7-459a-b688-68ef76fc615a
- Study
- Memory for AI Agents
- Research task
- 86ba16xrp
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "supermemory",
scenario: "memory-for-ai-agents"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Correct Application
Cognee✓ WorkedUses retrieved ACME history correctly by skipping the already-tried steps and moving to deeper SSO diagnostics instead of repeating the same troubleshooting.Hindsight✓ WorkedIt uses the stored ACME troubleshooting history to avoid repeating basic reset/cache/browser advice and moves to more advanced next-step investigation.Mem0✗ FailedKeeps applying old ACME SSO context after the rollout changed, so the later reply still falls back to SSO troubleshooting instead of the new email-password path.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com