Evidence · first-party tested/Best AI Tools for Memory for AI Agents

After ACME's rollout context was updated away from SSO, the UI still showed a '0 memory used' tag while the side panel continued to surface the old SSO memory at 71% and then 79% relevance on later replies, so it is not possible to confirm from the UI that the update reliably superseded the stale context.

⚠ Struggled🧾 artifact-verifiedinput + output shownTest date not recordedSupermemory
What was measured
Update and Correction Handling

Checks whether the tool can replace or invalidate outdated memories when corrected.

decisive for this rankingtransformation

A memory system has to replace or invalidate outdated information, or it will keep feeding the agent wrong context. (3 of 3 judges)

What was given, what came back

Test input: Client Relationship Memory · text · group: memory-for-ai-agents
Input — what we sent
The exact prompt
Session 1:
Use this under account_id: client_acme_001

ACME is a client using our AI support assistant. They prefer clear next steps and do not like repeated troubleshooting. Their team already tried password reset, clearing browser cache, and switching browsers. The issue is still happening only for users with SSO enabled.

Session 2:
Use this under account_id: client_acme_001

ACME came back today and said: "Our users still cannot log in with SSO. What should we try next?"

Draft a support reply that respects what they already tried and moves to the next useful step.

Session 3:
Use this under account_id: client_acme_001

Update the client memory: ACME is no longer using SSO for this rollout. They moved to email-password login for the first launch. Do not keep treating SSO as the active issue unless they mention it again.

Session 4:
Use this under account_id: client_acme_001

ACME says: "Some users are still unable to log in during launch testing."

Draft the next support reply.

Session 5:
Use this under account_id: client_beta_002

BetaCorp is a new client. They say: "Our users cannot log in for the first time."

Draft the first support reply for BetaCorp.

A client-support memory test that checks whether the assistant remembers prior troubleshooting, handles a changed rollout context, keeps client scope isolated, and avoids leaking one client’s history into another client’s support reply.

Why this input is hard
  • · client history recall
  • · avoid repeating troubleshooting
  • · update handling
  • · scope isolation between clients
  • · support response continuity
Output — unretouched
Output 1
Output 1
Output 2
Output 2
Provenance
Observation
296fccb1-080f-4f96-af1b-0af153818878
Evidence run
6e31afbb-34d7-459a-b688-68ef76fc615a
Study
Memory for AI Agents
Research task
86ba16xrp
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "supermemory",
  scenario: "memory-for-ai-agents"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on Update and Correction Handling
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com