The Q&A interface answered a natural-language question with a grounded response from the meeting record, including the specific date "6th August," and the report observed no hallucinations.

✓ Worked🧾 artifact-verifiedinput + output shownTest date not recordedNotta
What was measured
Chat with Notes / Ask Questions

Gives grounded answers with the cited moment and admits when unknown.

context, not decisivetransformation

Q&A over notes is useful, but it is an add-on to the capture and summarization job rather than a core measure of it. (3 of 3 judges)

What was given, what came back

Test input: AI Demos Daily Standup — 31 July 2026 · image · group: ai-meeting-notetaker
Input — what we sent
AI Demos Daily Standup — 31 July 2026
AI Demos Daily Standup — 31 July 2026

A real 25-minute technical engineering daily standup with 14 attendees and about 10 active speakers, used as the single parallel-capture meeting for evaluating AI meeting notetakers on transcription, diarization, summaries, action items, search/chat, and collaboration features.

Why this input is hard
  • · Transcription accuracy for real names, tool names, numbers, and technical jargon
  • · Speaker diarization across multiple active speakers
  • · Robustness to overlapping speech, crosstalk, and rapid turn-taking
  • · Join reliability for bot-based and botless capture
  • · Summary quality on identical source material
  • · Action-item extraction with correct owners and commitments
  • · Topic segmentation of standup updates
  • · Search and chat grounded in the meeting content
  • · Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage
Output — unretouched
image
Also checked on this input — same tool, 8 other criteria
Action-Item Extraction✓ WorkedThe action-item list extracted the real commitments from the call, formatted them as checkbox items with @mentions, and the report says owner assignment was correct with no false positives.Join Method & Reliability✓ WorkedThe bot-based Google Meet join was reliable in the tested call: Notta Bot appeared in the meeting list, admitted/managed normally, and the capture ran through the end of the session without disconnects or plan-limit cutoffs.Search Across Notes◐ MixedSearch works inside a meeting transcript through AI Chat and returns exact timestamps in plain text, but the timestamps are not clickable, and the report says this was not tested across meetings.Speaker Diarization◐ MixedSpeaker attribution was mostly correct, with nearly all speakers identified by name, but the transcript still showed some misattributed lines, so diarization was not fully reliable for every turn.Speaker Diarization⚠ StruggledThe transcript contained a line labeled with another notetaker’s name (HappyScribe), which indicates cross-tool contamination or labeling error and breaks speaker attribution for that segment.Summary Quality✓ WorkedThe generated meeting summary was comprehensive and skimmable, with structured sections such as Task & Issue Management and a mindmap-style organization that reflected the meeting flow.Topic Segmentation✓ WorkedThe meeting was segmented into useful topic blocks rather than one blob; the report names three sections, including Task & Issue Management, Individual Progress Updates, and API Benchmarking Task.Transcription Accuracy✓ WorkedNotta’s transcript capture was accurate on the evaluated standup: the report says it correctly captured names, tool names, numbers, and engineering jargon with no significant word-level errors, silent hallucinations, or misheard terms.
Provenance
Observation
b58378a0-5d37-4751-a711-f2d028fa4401
Evidence run
ace58582-3d1e-48ee-996c-9b3cd03f27a2
Study
AI Meeting Notetakers — Capture Accurate Transcripts, Summaries & Action Items From Live Calls
Research task
86baxegnv
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "notta",
  scenario: "ai-meeting-notetaker"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Chat with Notes / Ask Questions
Fathom✓ WorkedAsk Fathom answers direct factual questions from the meeting notes with grounded references; for one query it answered that a call was scheduled for 6th August and linked the supporting transcript mention.Fellow✓ WorkedAsk Fellow returned grounded answers to natural-language questions against the meeting notes, and the tested query produced a cited response rather than an unsupported hallucination.Fireflies.ai✓ WorkedAskFred answered a natural-language question with a specific grounded response ('August 6th') and relevant context, with no hallucination reported in the tested query.Granola✓ WorkedThe chat/Q&A surface gives grounded answers from the meeting record: on the tool-access question it says access was confirmed that day, cites both the notes and transcript, and identifies rerunning testing as the next step.HappyScribe✓ WorkedAI chat answers a meeting question with a grounded transcript-backed response, returning that the call was scheduled for '6th August' and explicitly indicating it is reading the transcription.MeetGeek◐ MixedIt answers direct grounded questions correctly, but the report records an incorrect answer on a speaker-dependent scheduling question, so chat is reliable for simple queries but weaker when attribution/context matters.Otter.ai✓ WorkedOtter’s AI Chat answered meeting questions with a grounded response and a specific timestamp, and the report says the answers were cited and free of hallucinations in the tested queries.
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com