-
Notifications
You must be signed in to change notification settings - Fork 0
How Questions Are Classified
Every question in Chiron carries three labels from three different literatures, and each option carries a fourth. This page explains what they mean and what they are for.
The short reason they exist: "this question is too easy" is an opinion, and opinions do not scale. A classification you can sort and count turns it into data. On one bank the tagging is what identified 51 recall-only questions as a group, so they could be removed as a group.
Bloom's revised taxonomy names the verb of a question. Six levels, roughly ascending:
Remember → Understand → Apply → Analyze → Evaluate → Create
Remember is retrieval: what does the rule say. Understand is interpretation: what does the rule mean. Apply puts it against facts it was not stated with. An exam that is all Remember tests whether you read the book; an exam that is mostly Apply tests whether you can practise.
Webb's DOK runs 1 to 4, and it is not a difficulty scale. A DOK 1 item can be brutally hard if the fact is obscure. What DOK measures is how much reasoning must happen between reading the question and knowing the answer.
- DOK 1 — recall. The answer is stated; you either have it or you don't.
- DOK 2 — a skill or concept. You must select, compare, or interpret before answering.
- DOK 3 — strategic thinking. Reasoning across a rule and a situation, often with justification.
- DOK 4 — extended thinking, across sources and time.
Bloom and DOK are not redundant. A question can be Understand at DOK 1 (recall of a definition, lightly rephrased) or Understand at DOK 2 (you must interpret a preamble against three near-neighbours). The pair pins down what the item really asks.
SOLO — the Structure of the Observed Learning Outcome, from Biggs and Collis — describes how organised a response is:
| Level | The answer shows |
|---|---|
| Prestructural | Nothing relevant. A miss. |
| Unistructural | One relevant element, in isolation |
| Multistructural | Several relevant elements, unconnected |
| Relational | The elements connected into a coherent whole |
| Extended abstract | The whole generalised beyond the material given |
Chiron's distinctive choice is to apply SOLO to the distractors — the wrong answers.
That is the part worth slowing down for. In an ordinary question bank a wrong answer is just wrong, and choosing it tells you nothing except that you failed. If instead each wrong answer is built to embody a particular structure of misunderstanding, then choosing it says something specific about what you believe.
A worked example. A question asks what distinguishes the Council of Europe's Framework Convention on AI from other AI governance instruments. The correct answer is that it exists to keep AI activity consistent with human rights, democracy and the rule of law. One wrong answer offers harmonised technical product-safety requirements and conformity assessment for high-risk systems — and is labelled relational.
Relational, not random. Whoever picks it has organised the material well: they know product-safety conformity assessment is a real regime with real machinery. They have simply attached it to the wrong instrument — that is the EU AI Act. A learner who chooses it does not need to be told the answer. They need to be told which two things they have merged.
The other distractors on that question do similar work: one offers a voluntary risk-management framework (the NIST AI RMF's map-measure-manage shape), another a certifiable management-system standard (ISO/IEC 42001). Each is a named, plausible neighbour rather than filler.
A distractor that nobody would ever choose teaches nothing. SOLO levels are how a reviewer judges whether a wrong answer is doing its job.
Before a question is drafted, its purpose is declared in a sentence: what this item is meant to establish about the candidate.
This is a forcing function, not documentation. A generator asked for "a question about the Framework Convention" produces something about the Framework Convention. A generator asked to test whether a candidate can distinguish that Convention's objective from adjacent instruments has to do something harder, and the distractors follow from the goal rather than being invented afterwards.
The goal statement is also what a human reviewer checks the question against. "Does this item do what it says it does" is a far more answerable question than "is this a good item."
The labels are not decoration on the review screen. They are the grounds for a decision:
- The item claims DOK 2 but can be answered by matching a phrase from the stem to a phrase in an option → the tag is wrong, and the item needs revising rather than approving.
- A distractor is labelled relational but is obviously absurd → it will never be chosen, so it is doing no diagnostic work.
- The goal says the item distinguishes two instruments, but the distractors are not the neighbouring instruments → the item does not do what it claims.
- The bank is overwhelmingly Remember / DOK 1 → the domain is teaching recognition, not practice, whatever any individual question looks like.
Bloom's revised taxonomy (Anderson and Krathwohl, 2001, revising Bloom 1956); Webb's Depth of Knowledge (Norman Webb, 1997); the SOLO taxonomy (Biggs and Collis, 1982). All three are long-established in assessment design. Chiron's contribution is not the taxonomies — it is applying all three at once, at generation time, and putting SOLO on the distractors.
Start here
| You are… | Go to |
|---|---|
| Sitting an exam | Studying with Chiron |
| Sceptical of AI questions | Can You Trust These Questions? |
| Here for the learning science | How Questions Are Classified |
| A technical reader | How Chiron Works |
| Reviewing security or privacy | Security · Privacy |
The method
The exams
Boundaries
Chiron is the teacher. Antron is the cave he taught in.