Session Flo logoSession Flo
11 min readInteractivity

Confidence Wagers: Making Quizzes Sharper

Ask people how sure they are and a quiz stops measuring guesswork. Here is how to score confidence wagers, when they sharpen a session, and when they wreck it.

By Session Flo

Key takeaways

  • Confidence wagers ask for two things per question - the answer and how sure you are - and the second one is where the teaching value sits.
  • Confident and wrong is the state worth finding; a plain quiz cannot distinguish it from a lucky guess.
  • Three confidence levels are enough. Five gives you more precision than a room can use in ninety seconds.
  • Penalise high-confidence wrong answers, or the wager is free and everyone maxes it out by round three.
  • Cap or ban all-or-nothing final rounds - they overwrite everything the earlier rounds measured.
  • Debrief the confident-wrong questions first and the everyone-got-it questions never.

Confidence wagers are a quiz format where each person submits an answer and a stake: how sure they are, expressed as a level or a number of points. Get it right with high confidence and you score well; get it wrong with high confidence and you lose more than a cautious answer would have cost. The scoring is the point, because it forces people to separate what they know from what they are guessing.

A standard quiz gives you one bit of information per person per question: right or wrong. That number hides the difference between someone who knows the answer cold and someone who narrowed it to two options and flipped a coin. In a training context those are opposite situations that need opposite responses, and you cannot tell them apart from a bar chart.

This piece covers the scoring rules that behave properly, how to run wagering live without slowing the quiz to a crawl, what the results actually tell you, and the four ways the format goes wrong - including the last-round all-in that undoes an otherwise good session.

What the wager adds that a plain score does not

Split any group's answers into four boxes: right and sure, right and unsure, wrong and unsure, wrong and sure. A plain quiz merges the first two and the second two. The four-box view is far more useful, because each box needs a different intervention.

Right and sure needs nothing - move on. Right and unsure is knowledge that exists but is not trusted, and it usually needs a confirmation and a reason rather than re-teaching. Wrong and unsure is an honest gap: teach it. Wrong and sure is the dangerous box, because that is where people act on the belief without checking, and it is almost always the item that is worth the debrief time.

Retrieval practice already makes quizzes a better learning tool than re-reading the material. Adding the wager makes the retrieval more deliberate: you cannot stake your points without asking yourself how you know what you think you know, and that question is doing most of the work.

Plain right/wrong quiz

  • One bit of information per answer
  • Lucky guesses look identical to solid knowledge
  • Debrief time gets spent on whatever scored lowest
  • Fast answerers do well regardless of certainty
  • No signal about which misconceptions are held firmly

Quiz with confidence wagers

  • Two signals per answer: correctness and certainty
  • Guesses declare themselves as low-stake answers
  • Debrief targets the confidently-wrong items first
  • Reflection is built into every question
  • You leave with a list of misconceptions worth correcting

Scoring schemes that behave

The whole format lives or dies on the maths. If a high stake costs nothing when you are wrong, staking high is always correct and everyone works that out within three questions. The wager has to be a genuine trade-off.

Two schemes cover almost every situation. The three-level scheme is the one to use by default in live sessions: low confidence scores 1 for a correct answer and 0 for a wrong one, medium scores 2 or minus 1, high scores 3 or minus 2. It is easy to explain in fifteen seconds, easy to do in your head, and the negative on high confidence is enough to make people think.

The points-budget scheme suits longer quizzes and competitive settings: each person gets a pot of, say, 20 points to spread across 10 questions, staking between 0 and 5 on each. Correct answers win the stake, wrong ones lose it. It is more engaging and much more prone to gaming, so it needs a per-question cap and a rule about the final round.

The three-level scheme in practice

Announce the numbers before question one and put them on screen. 'Sure is worth three but costs you two if you are wrong. Fairly sure is two and minus one. Guessing is one and zero - guessing is free.' The last sentence matters more than the others: making low confidence costless is what stops people bluffing to save face.

Do not change the numbers mid-quiz. People calibrate to the scheme they were told, and shifting it invalidates the comparison between early and late rounds - which is exactly the comparison you want if you are running the same quiz before and after training.

Low confidence: +1 / 0

Guessing is free. This is the pressure valve that keeps the data honest, and the reason people are willing to admit uncertainty in front of colleagues.

Medium confidence: +2 / -1

The default choice for most answers. A small penalty makes people pause without making the round feel punitive.

High confidence: +3 / -2

Reserved for knowledge people would act on without checking. The penalty is what makes a high stake mean something.

Per-question cap

In a points-budget version, no stake above a quarter of the pot. Without it, one lucky final answer decides the whole quiz.

Running it live without killing the pace

The obvious risk is time. Two decisions per question instead of one can turn a snappy ten-question quiz into a twenty-minute slog, and a slow quiz loses the room regardless of how good the data is.

Keep both inputs on one screen so the answer and the stake are a single submission. Allow 30 to 45 seconds per question for factual recall and up to 60 for anything that requires reading a scenario. Do not add extra time for the wager - people who need longer to choose a stake are usually reconsidering the answer, and the time pressure produces more honest calibration than a leisurely re-think does.

Ten questions is the ceiling for one sitting. Confidence judgements are a second cognitive task on top of the recall, and the extra load shows up as sloppier staking somewhere around question twelve. If you have thirty questions of material, run three rounds of ten across a training day rather than one long block.

A ten-question round, timed

1

Explain the scoring (60 seconds)

Three levels, the numbers on screen, and the line 'guessing is free'. Say explicitly that the aim is accurate self-assessment, not maximum points.

2

Run a throwaway practice question (45 seconds)

Use something everyone knows or nobody could know, so the first real question is not also the first time anyone has staked points. Never score the practice question.

3

Ten questions at 30-45 seconds (7 minutes)

Reveal the correct answer after each question, but hold the confidence breakdown back until the end so nobody adjusts their staking to match the room.

4

Show the four-box split (2 minutes)

For each question, what share of the room was right-and-sure versus wrong-and-sure. This is the output you came for, and it is worth putting on screen in full.

5

Debrief two questions (6 minutes)

Take the two with the highest confidently-wrong share. Ask what made the wrong answer feel obviously right - that reasoning is the misconception, and naming it out loud is what corrects it.

6

Close with a calibration line (1 minute)

'Across the room, high-confidence answers were right about four times in five.' State it as the room's own number, then say what you will re-test next month.

Numbers to design around

These are design guidelines rather than research findings: they are the settings that hold a room's attention while still producing usable calibration data. Treat them as a starting point and adjust once you have seen how your group stakes.

3
Confidence levels - more precision than a room can use
30-45 sec
Per question, answer and stake together
10
Questions maximum in one wagering round
2
Questions to debrief properly, chosen by confident-wrong share

The four failure modes

1

Everyone stakes high on everything

Your penalty is too small or absent. Increase the loss on high confidence until it hurts, and remind the room that a low-confidence correct answer still scores. If it persists, the group has read the exercise as a competition rather than a diagnostic - which is a framing problem, not a scoring one.

2

Everyone stakes low on everything

Usually a safety issue: people do not want to be publicly wrong with confidence. Run the wagering anonymously, show only aggregate confidence, and never put a name next to a confidently-wrong answer on screen. Risk-taking in front of colleagues is not evenly distributed, and a format that rewards it will systematically flatter the people who are comfortable being loud.

3

The final round decides everything

Any all-in last question turns nine rounds of careful calibration into a coin flip. Either cap the final stake at the same level as every other question, or make the final round unscored and use it purely for discussion.

4

The debrief goes to the lowest-scoring question

Tempting, but wrong. A question everybody got wrong with low confidence is a known gap and needs teaching, not discussion. Spend the scarce debrief minutes on the question people were sure about and wrong about, because that is the belief that will otherwise leave the room intact.

Where confidence wagers help and where they hurt

The short rule: use wagers when being wrong has consequences outside the room, and skip them when the quiz exists to make a Friday afternoon better. A social quiz with confidence staking is a slower social quiz.

They are strongest in safety, clinical, legal, security and technical onboarding contexts, and in any refresher where the group believes it already knows the material. The confident-wrong count in those sessions is usually the most persuasive slide in the whole programme.

Pros

  • Separates solid knowledge from lucky guessing on every single item
  • Surfaces confidently-held misconceptions, which are the ones that cause real errors
  • Builds a habit of self-assessment that carries beyond the quiz
  • Gives you a defensible before-and-after measure for a training programme
  • Makes a ten-question quiz meaningfully harder without adding harder questions

Cons

  • Adds a second decision per question and slows the pace
  • Rewards comfort with risk, which is unevenly distributed across a group
  • Negative scoring can feel punitive in a compliance or assessment context
  • Gameable in a points-budget format unless you cap stakes
  • Pointless for pure fun quizzes, where the extra thinking is just friction

Setting it up

Mechanically you need a quiz that accepts two inputs per question and a scoring engine that applies asymmetric points. If your tool only does right and wrong, you can fake it with an extra question - answer first, then 'how sure were you?' - but you will pay for it in pace, and the second question invites people to revise their certainty after the fact.

Session Flo handles the two inputs as one submission and applies the wager rules automatically, so the scoreboard and the confidence breakdown come out of the same activity without you doing arithmetic in front of the room. Whatever you use, check that participants can see the scoring rule while they answer - people calibrate badly against a rule they half-remember from a slide two minutes ago.

Then re-run it. The number worth tracking is not the score, it is the gap between how often people staked high and how often high stakes paid off. Run the same ten questions six weeks later and a shrinking gap tells you the training changed something, which is more than a rising average score ever proves.

Frequently asked questions

How many confidence levels should a quiz use?

Three. Low, medium and high map cleanly onto guessing, fairly sure and would-act-on-it, and people can pick between them in a couple of seconds. Five levels sound more precise but the middle three collapse into 'not sure' in practice, and the extra choice costs you time on every question. A continuous 0-100 slider is worse still in a live room, though it is fine for asynchronous assessment.

Should wrong high-confidence answers lose points?

Yes, or the wager means nothing. Without a penalty, staking high is strictly optimal and the room works that out within three questions, at which point you have added a click per question and learned nothing. Minus two against a plus three is enough. In a formal assessment context, be transparent about the negative marking before anyone starts, and consider a floor of zero for the round.

Do confidence wagers disadvantage cautious people?

They can, which is why the low-confidence option must score positively for a correct answer and why the results should be anonymous by default. If the leaderboard rewards boldness rather than accuracy, cautious but knowledgeable people finish mid-table and quietly disengage. Reporting calibration - how often high stakes were correct - alongside raw score fixes most of it.

How long does a wagering quiz take?

About seventeen minutes for a proper round: a minute to explain scoring, a practice question, ten scored questions at 30 to 45 seconds each, two minutes on the confidence breakdown and six on debriefing two questions. If you have less than fifteen minutes, cut questions rather than the debrief, because the debrief is where the format pays for itself.

Can you use confidence wagers with team scoring?

Yes, and it changes the activity for the better. Teams have to agree a stake as well as an answer, and that argument is where the learning surfaces - someone has to say out loud why they are sure. Give teams 60 to 90 seconds per question to allow the discussion, and keep the round to six or seven questions.

What do you do with the results afterwards?

Publish the four-box split per question in the session recap, without names. List the two or three items with the highest confidently-wrong share as the things to re-check, and schedule a repeat of the same questions four to six weeks later. Comparing the two confident-wrong counts is a far better measure of whether the training worked than the average score.

Keep reading

Related articles

Ready to run better sessions?

Create interactive events with live polls, quizzes, icebreakers, and more.