Observability is weaker here: Sessions 1–3 expose only short acknowledgments, and the detailed Session 4 handoff note still requires re-verification against the project's own stored sessions.
What was measured
Observability and Debugging
Checks whether developers can inspect what happened, including stored and retrieved memories.
context, not decisivecapability
Inspecting stored and retrieved memories helps teams debug and trust the system, but it does not itself measure memory performance. (3 of 3 judges)
What was given, what came back
Test input: Team Handoff / Project Continuity Memory · text · group: memory-for-ai-agents
Input — what we sent
The exact prompt
Session 1: Use this under project_id: ai_demos_memory_use_case We are working on an AI Demos use case called Memory for AI Agents. The goal is to help users understand which memory tools are actually useful for real agent workflows. We are not promoting any tool. We are testing whether memory can help with real continuity: personal work brain, client relationship memory, and team handoff. Session 2: Use this under project_id: ai_demos_memory_use_case Important project rules: - No observation without proof. - Screenshots and artifacts are primary evidence. - Inputs must help rank tools, not just prove that tools can store one fact. - Memory should be checked for retrieval, update handling, scope control, deletion or retirement, and observability. - The page should stay practical and user-facing, not only technical. Session 3: Use this under project_id: ai_demos_memory_use_case Project direction changed slightly. The old input set was too QA-style and not relatable enough. The new direction is to use real workflows: personal work brain memory, client relationship memory, and team handoff/project continuity memory. Session 4: Use this under project_id: ai_demos_memory_use_case I am unavailable tomorrow. Create a handoff note for an intern who needs to continue this use case. The note should explain: 1. What this use case is about. 2. What the current testing direction is. 3. What rules they must follow before writing observations. 4. What artifacts they need to capture while testing. Session 5: Use this under project_id: unrelated_sales_agent_project We are building a sales email agent for a different project. Create a short kickoff note for the team.
A project continuity and handoff test where the assistant must remember project goals, project rules, and a changed testing direction, then produce a useful handoff note for an unavailable team member without leaking context into an unrelated project.
Why this input is hard
- · project-level memory
- · decision and rule retention
- · changed direction handling
- · handoff continuity
- · scope separation across projects
Output — unretouched





Also checked on this input — same tool, 3 other criteria
Memory Capture Quality◐ MixedStores project context, rules, and direction changes under the project container, but the visible replies were only generic acknowledgments rather than a visible restatement of the stored content.Relevant Retrieval◐ MixedCannot confidently retrieve the right project memory for the handoff task: the detailed Session 4 note does not clearly match the project's own stored sessions, so the result needs re-verification.Scope Control✓ WorkedKeeps an unrelated `unrelated_sales_agent_project` container separate from the AI Demos project, with the kickoff note staying on the sales-agent topic and not inheriting AI Demos or BetaCorp context.
Provenance
- Observation
- 86d154fe-328f-4f54-b64a-879ce687540e
- Evidence run
- 6e31afbb-34d7-459a-b688-68ef76fc615a
- Study
- Memory for AI Agents
- Research task
- 86ba16xrp
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "cognee",
scenario: "memory-for-ai-agents"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Observability and Debugging
No other tool was measured on this criterion for this input.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com