AI Demos Research — the structured-intelligence platform. Every verdict on these pages opens to the execution behind it.
Tools/Freshdesk Freddy/Conversation activity can be inspected·AI Customer Support Chatbots benchmark·Graded 12 September 2026

Can Freshdesk Freddy show enough conversation activity to inspect what happened in a thread?

Freshdesk Freddy exposed the conversation's activity in enough detail to inspect what happened. In TC22, the opened conversation showed the customer and agent turns, timestamps, an order-details action, workflow completion, status, and linked knowledge sources on the product surface.

1 of 1 test case passed

Every test case this benchmark pins to the scenario has an accepted result.

Pass rate100%1 of 1 gradable
Coverage1 of 1pinned test cases graded
1PassEvery expectation held against the reference.
0FailNothing was shown not to hold either.
0Not gradableEvery result here could be graded.
0UntestedEvery test case in this scenario has a result.
The pass rate is a summary. The one row below is the evidence — each opens onto the stimulus as sent, the expectations it was checked against, what the tool returned, and the proof.

The test case

One result per test case: Pass, Fail, Not gradable, or Untested. Each is graded on its own against the expectations registered with its test-case version — the scenario result is just this row counted.

graded against the test-case expectations · unit of grading = test case
TC22Find the earlier ⟨shipped-order⟩ conversationPassTicket logs found and opened Conversation #14. The transcript then showed order 41982 with Status: Shipped.Evidence
The stimulus, as sent

Exact stimulus wording was not preserved for this run.

What the tool returned

The tool's reply was not preserved in the evidence for this run.

Proof
Screen recording of Ticket logs search opening Conversation #14 and showing the shipped-order turn.
cited at▶ 0:01▶ 0:11▶ 0:31▶ 0:46
Graded against
expected_behaviorassertsThe conversation can be found and openedThe conversation was found in Ticket logs and opened.
Why this result

Ticket logs found and opened Conversation #14. The transcript then showed order 41982 with Status: Shipped.

Also observed on this row

Search by ticket id. The Ticket logs search box found the conversation using ticket id 14. The opened log matched the shipped-order conversation.

Workflow trace visibility. Conversation logs show per-answer knowledge-source attribution and workflow traces. The transcript includes the order lookup action, the generated response, and workflow completion.

Run · tested by Ajay Vekhande · evidence submitted by 4 September 2026
Test case · TC22 v1

Rows are never hidden behind the percentage, and never collapsed into one another. A test case with more than one run keeps the latest gradable result here and moves the earlier ones into the history at the foot of the page.

Also observed

Things recorded that no expectation in this scenario covers. They do not move the result, but they are real.

recorded, not graded
·The Ticket logs search box found the conversation using ticket id 14. The opened log matched the shipped-order conversation.
·Conversation logs show per-answer knowledge-source attribution and workflow traces. The transcript includes the order lookup action, the generated response, and workflow completion.

How this scenario is graded

The rule that produced every row above, printed rather than described. It is the same rule for every tool tested on this scenario.

The unit of grading

The test case. One result per test case per tool: Pass, Fail, Not gradable. A pinned test case with no accepted result reads Untested. No Partial.

The rules
  • Pass — every expectation on the test-case version holds against the registered reference, and nothing in the reply contradicts the reference.
  • Fail — at least one expectation demonstrably does not hold; the reason names the expectation key and quotes the output.
  • Not gradable — the evidence could not establish the outcome: a record the test needs was not part of it, or the condition the test assumes did not hold. Never inferred as a fail; the row says what could not be established.
Note on this scenario

No scenario rubric is pinned; each verdict was graded against the expectations registered on the test-case version (V1, 2026-09-10).

Configuration and setup

The software, surface and connected system this run used, and when it was tested.

Software that produced the output
Freshdesk Freddy
Surface
Ticket logs screen
Tested
By 4 September 2026 · Ajay Vekhande

Other tools on this scenario

13 products are in this benchmark. Tiles change on their own as results are published.

Open the scenario page →
Adauntested · 0/1
Botpressuntested · 0/1
Chatbaseuntested · 0/1
CustomGPT.aiuntested · 0/1
Decagonuntested · 0/1
Freshdesk Freddy1 pass · 0 fail · 100% · 1/1
FS Agent (DIY control)untested · 0/1
Gorgiasuntested · 0/1
Help Scoutuntested · 0/1
Intercom Finuntested · 0/1
Kommunicate1 pass · 0 fail · 100% · 1/1
Tidio Lyrountested · 0/1
Zendesk AIuntested · 0/1

Freshdesk Freddy on the other scenarios

26 scenarios in this benchmark.

Open the tool page →
A configured rule requires escalationuntested · 0/1
Action requires confirmation1 pass · 0 fail · 100% · 1/1
Agent actions/tool calls can be inspecteduntested · 0/1
Agent cannot resolve the requestuntested · 0/1
Answer is available in the knowledge baseuntested · 0/3
Answer is not available in the knowledge baseuntested · 0/3
Configured instruction changes agent behavioruntested · 0/1
Configured restriction is respecteduntested · 0/1
Conversation activity can be inspected1 pass · 0 fail · 100% · 1/1
Correction changes future behavior where the product claims learninguntested · 0/1
Customer asks in different words than the knowledge base uses1 pass · 0 fail · 100% · 1/1
Customer communicates in another supported languageuntested · 0/1
Customer corrects information given earlier1 pass · 0 fail · 100% · 1/1
Customer switches languages during the conversationuntested · 0/1
Earlier conversation/session needs to be continueduntested · 0/1
Feedback can be submitteduntested · 0/1
Information from earlier in the conversation is needed later0 pass · 1 fail · 0% · 1/1
One message contains multiple requestsuntested · 0/1
Reported analytics match what actually happeneduntested · 0/1
Requested action cannot be completed1 pass · 0 fail · 100% · 1/1
Request is unclear and needs clarification0 pass · 1 fail · 0% · 1/1
Request requires choosing the correct source or action1 pass · 0 fail · 100% · 1/1
Retrieve information from an external system1 pass · 0 fail · 100% · 1/1
Update something in an external systemuntested · 0/1
User explicitly asks for a humanuntested · 0/1
User is not authorized to perform the actionuntested · 0/1

Where this sits in the benchmark

This page is one cell of a larger study: one tool, one scenario. Only this benchmark's frame appears here.

LevelNameScope
BenchmarkAI Customer Support Chatbotsv1 · 26 scenarios · 13 products · not yet frozen
CapabilityAnalytics and observabilityC8
ScenarioConversation activity can be inspectedS22 · weight 1.0 · role context
Rubricnone pinnedgraded against the test-case expectations
Test casesTC221 pinned
ToolFreshdesk Freddytool

History of this result

What has happened to this result since it was first published. Runs and grades are never overwritten: a retest or a re-grade publishes a new result and keeps the earlier one readable.

from the publication record
22 September 2026First publishedAI Customer Support Chatbots v1

Act on this result

Nothing filed here edits the run or the grade. A challenge opens a review, and a review can produce a new run or a re-grade — which becomes the current result and leaves this one in the history.

signals are counted, not scored
This matches what I see

You run the same kind of test against your own setup and get the same behaviour.

Agree
This does not match

Yours behaves differently. Tell us what you got, with a screenshot if you have one.

Disagree
Point out an issue

Something here is wrong — a reference value, a transcription, a grade.

Report an issue
Request a retest

On a newer build, a larger dataset, or your own setup.

Request a retest
We have fixed this

Tell us what changed and we schedule a rerun of the failing test case. The old result stays as history.

Vendor notice
Filed against this evidenceNothing yet. Challenges, counter-evidence and fix notices appear here with their outcome, and stay on the page after they are resolved.
Cite this result
aidemos.com/benchmarks/ai-customer-support-chatbots/results/freshdesk/conversation-activity-can-be-inspected · 1 pass · 0 fail · coverage 1/1 · graded 2026-09-12

The same record is available as structured data through the AI Demos MCP server, with the counts, the coverage and every per-test-case reason carried as fields.

Provenance and identifiers
toolid edaa8b9a-3e60-4039-a50b-6e3ab32b99b6 · slug freshdesk
proofsbytes 1860994 · sha256 12687f8db0c201b4e80ee8d71c18291bdfc08f058058563b48fde74ef8b13879 · filename Freshdesk_C8_S22_TC22_AnalyticsObservability_InspectConversationActivity.mp4 · uploaded_at 2026-09-04T12:21:28.166000+00:00
reviewreviewer model claude-opus-5 · runtime Claude Code · methodology_version v1-milestone-1 · reviewed_at 2026-09-12T05:04:01.670467+00:00 · proposal_sha256 5d863f2b31f5f8393fda2599d5a28de102b5955025735739fe790dc9335e6843 · submission_sha256 a2ab58958643b246e158895a5ad5ffa4108f680da295d31c7479443d67df8308
rubric
storagepass → worked · fail → failed · not gradable → not_gradable
scenarioid 327514aa-4971-4873-bce9-60ebfc7ae0d5 · code S22 · slug conversation-activity-can-be-inspected
benchmarkslug ai-customer-support-chatbots · version 1 · version_id 1bb1d77b-2b3e-4c25-ba37-3c6f4d6cfd54 · pinned_digest 18e7943aa11dc5c59f6d1da64a57ff29816f6c8baed6705cdd7a57861273769a
researchername Ajay Vekhande · round 1
test case versions pinnedcode TC22 · state being_written · version_id 17ba92f3-9820-43e8-aade-70d0e0beb7ea