Across the reviewed quiz sessions, each question explanation broke down the wrong options individually instead of only stating the correct answer, giving per-option elimination reasoning.

✓ Worked🧾 artifact-verifiedoutput onlyTest date not recordedPDF to Quiz
What was measured
Are the Questions Actually Good?

Whether the questions test understanding at the requested difficulty level rather than only surface-level recall or low-quality near-duplicates.

decisive for this rankingtransformation

Question quality is the heart of the job here; weak or shallow questions mean the tool is failing at generating useful quizzes. (3 of 3 judges)

What was given, what came back

Input — what we sent
No input — this is a capability finding
The observation is about the tool itself rather than one test input, so there is nothing to show on this side by design.
Provenance
Observation
f0f58995-e7b7-4fdb-bb6e-424a1361d956
Evidence run
de7b348e-4326-4fbf-87c1-b104750bc459
Study
Generate Quizzes and Practice Tests from Study Material Using AI
Research task
86ba51dce
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
output only
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "pdf-to-quiz"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Are the Questions Actually Good?

No other tool was measured on this criterion for this input.

Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com