Exposes recall provenance directly in the UI: the Last Recall Response panel shows the source graph completion, evidence chunks, dataset ID, and document/chunk IDs.
What was measured
Observability and Debugging
Checks whether developers can inspect what happened, including stored and retrieved memories.
context, not decisivecapability
Inspecting stored and retrieved memories helps teams debug and trust the system, but it does not itself measure memory performance. (3 of 3 judges)
What was given, what came back
Test input: Personal Work Brain Memory · text · group: memory-for-ai-agents
Input — what we sent
The exact prompt
Session 1: Use this under user_id: founder_001 I run a small AI product/research team. When you help me, remember how I work: - Keep outputs short, direct, and copy-paste ready. - Do not make writing sound too polished or motivational. - Always mention what proof or artifact is needed before making a strong claim. - If a task is risky or unclear, tell me the safest next step instead of guessing. Session 2: Use this under user_id: founder_001 Today I am testing tools for an AI memory use case. I want to show users that memory is not just "remember my favorite color." It should help an assistant continue real work across days, remember my working style, and avoid repeating the same explanation again. Create a short internal update for my team about what I worked on today and what we should test next. Session 3: Use this under user_id: founder_001 Now write a formal email to a potential enterprise partner asking if they are open to a product demo next week. Keep it professional.
A multi-session personal assistant memory test where the user first sets working-style preferences, then asks for an internal update, and finally requests a formal partner email to check whether the assistant applies memory selectively and appropriately across different writing tasks.
Why this input is hard
- · work-style preference memory
- · cross-session retrieval
- · tone adaptation by task
- · proof-first behavior
- · avoiding overgeneralization of memory
Output — unretouched

Also checked on this input — same tool, 3 other criteria
Correct Application✓ WorkedApplies the stored working style to a new internal update by keeping the reply short and proof-oriented while blending in the current memory-testing project context.Memory Capture Quality✓ WorkedStores a compact working-style profile as durable graph-backed memory: concise/direct output, no over-polished tone, proof-before-claims, and safest-next-step behavior for risky tasks.Relevant Retrieval✓ WorkedRecalls prior working-style memory in later sessions through GRAPH_COMPLETION, exposing evidence chunks plus dataset and document/chunk IDs rather than a simple memory-hit flag.
Provenance
- Observation
- 22461c92-39c7-4585-a0b6-7210ea110012
- Evidence run
- 6e31afbb-34d7-459a-b688-68ef76fc615a
- Study
- Memory for AI Agents
- Research task
- 86ba16xrp
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "cognee",
scenario: "memory-for-ai-agents"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Observability and Debugging
No other tool was measured on this criterion for this input.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com