Quiz Generation
This benchmark evaluates how AI tools turn supplied study material into quizzes and practice tests, then grade, explain, and track what learners actually understood.
What this benchmark is
This benchmark is for AI products that work from supplied study material. It asks whether a tool can turn that material into quizzes and practice tests, then deliver them, grade responses, explain mistakes, and track understanding.
Buyers can use it to compare products for classroom and learner use. The benchmark shows whether the tool handles assessment content, answer evaluation, learner feedback, ongoing practice, and reporting in ways that stay faithful to the source material.
Version 1 uses the approved Electric Circuits textbook corpus. The registered resource Quiz generation fixture material — Foundation Science 8 chapters supplies those chapter materials for this benchmark.
In scope
- Turning supplied study material into quizzes and practice tests.
- Delivering quizzes to individual learners or classes.
- Grading fixed-answer, free-form, and partly correct responses.
- Giving feedback that explains what is correct, missing, or wrong.
- Adapting later practice to strengths and weaknesses, including across sessions.
- Showing learner progress over time.
- Showing class reports about what the class struggled with.
Out of scope
- Product attributes, which are captured separately and never scored.
Capabilities included
The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.
| Capability | What it means here | Scenarios |
|---|---|---|
| Quiz Generation | Creates valid quizzes from supplied material that test the intended knowledge and include sufficient information for answers to be evaluated. | 5 |
| Quiz Delivery | Lets learners take and submit quizzes in the intended individual or class setting while recording each attempt correctly. | 2 |
| Grading | Determines how correctly a learner answered, including fixed-answer scoring, semantically equivalent wording, and partial understanding. | 3 |
| Answer Feedback | Explains what the learner understood correctly and what remains wrong, missing, or incomplete. | 2 |
| Personalized Practice | Changes later practice according to the learner's demonstrated strengths and weaknesses, including across sessions. | 2 |
| Learner Progress | Shows learners how their performance changes over time and which topics still need work. | 1 |
| Class Reports | Shows teachers how learners performed and which questions or topics the class struggled with. | 1 |
Scenarios included
A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.
Quiz Generation
5 scenariosGrading
3 scenariosAnswer Feedback
2 scenariosPersonalized Practice
2 scenariosHow the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Scope includes quizzes and practice tests
The benchmark covers quizzes and practice tests from supplied study material, then delivering, grading, explaining, and tracking them.
Open rubrics and scoring
Rubrics, weights, scoring, and ranking methodology remain open and unpinned.
Product attributes are separate
Product attributes are captured separately and never scored.
Class reports show struggle, not trend
Class reports cover what the class struggled with; do not invent an over-time class-performance trend.
Test case implementation and fixture design remain the open stage.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
Quiz generation fixture material — Foundation Science 8 chapters
Supplies the approved Electric Circuits chapter materials used as the benchmark corpus.