Question quality benchmark
81.3 on a held-out eval.
Field average: 67.9.
Most AI quiz tools never publish how their questions score. This is our number, the rubric behind it, and the instructions to run the same eval on your own deck.
Used by 1M+ students on Jungle.
Free tier, no credit card.
Leaderboard · 3 documents · 4 criteria
Measured 2026-04-24 · Jungle internal Quality Comparison panel · same three documents, same four criteria, same scoring sheet for every tool.
How we scored it
A number you can audit, not a number you have to trust.
Three held-out documents
Three source documents the generator was never tuned on: a lecture slide deck, a textbook PDF, and a third doc. Held-out means no tool got to practice on them first.
0 / 3 / 7 / 10 anchors
Each criterion is graded on a four-point anchor scale (broken / weak / solid / exemplary), not a 1-5 Likert. The 0-40 raw score is renormalized to 0-100. That renormalized number is what you see on the board.
Four equal-weight criteria
Factual correctness, clarity, and distractor quality are scored per card; type coverage once per deck. No criterion is weighted higher than another, so a tool can't buy its score back with a single strength.
What it does not claim
It is one snapshot on three documents, measured 2026-04-24 on Jungle's internal Quality Comparison panel. It is not a license-exam pass rate and not a promise that every card is perfect. It is a floor: how often each tool hands you a card worth studying from.
What the 23.5-point spread means
In a 200-question deck, a tool scoring 57.8 hands you roughly forty cards that are wrong, ambiguous, or test-takeable on length tells. A tool scoring 81.3 hands you maybe twelve. You still spot-check, but the floor is higher.
Criteria
Factual correctness
Every question's correct answer is grounded in the source document. We verify against the actual PDF / slide content, not the model's pretrained knowledge. No source span, the card scores 0.
Clarity
The stem is unambiguous. A well-prepared student should be able to identify which option matches without re-reading the question three times.
Distractor quality
Wrong answers are plausible and similar in length. No 'all of the above', no obviously wrong throwaway options. The question rewards understanding, not test-taking heuristics.
Question type coverage
A good deck mixes recall, application, comparison, and case-based questions. Scored once across the whole deck, not per card, so a wall of single-best-answer MCQs is penalized.
Distractor quality, made concrete
The criterion that splits the leaderboard hardest is distractor quality. Here is one card Studyly wrote from a cardiology deck. Every wrong option is a real valve lesion, length-matched, defensible to a classmate. That is what an 81.3 looks like one card at a time.
From a cardiology deck · MCQ
Q1 / 200
A systolic murmur is loudest at the apex and radiates to the axilla. Which valve lesion does this describe?
Pick one. This is one of two hundred Studyly wrote from the deck.
See for yourself.
Drop your next lecture deck and read the first ten questions against the same four criteria. Free tier, no credit card.