Session Flo logoSession Flo
13 min readResearch

Wisdom of the Crowd: When Group Estimates Beat Experts

Averaging a group can beat the best individual in the room — but only when the estimates are independent. Most meetings destroy that independence in the first thirty seconds.

By Session Flo

Key takeaways

  • Aggregated group judgement beats the average member and often the best member, but only when estimates are made independently before anyone speaks.
  • The moment a senior person says a number out loud, the crowd stops being a crowd and becomes an echo of that number.
  • Collect silently, reveal all at once, and use the median rather than the mean when the spread has long tails.
  • Eight to twelve independent estimators is enough for most workplace questions; beyond about twenty the accuracy gain flattens.
  • The spread matters more than the central number: a wide disagreement means people are answering different questions, not that they are bad at estimating.
  • Use a specialist, not a crowd, for questions requiring specific technical knowledge that most of the room simply does not have.

The wisdom of the crowd is the finding that aggregating many independent judgements often produces an estimate closer to the truth than any single member's guess, including the expert's. Individual errors scatter in both directions and largely cancel; what survives the averaging is the signal everyone partly had.

The condition doing all the work in that sentence is independent. Errors only cancel if they are not correlated, and nothing correlates a room's errors faster than hearing what somebody else thinks first. This is why the effect is robust in research and almost absent in the average planning meeting.

This article covers the four conditions the effect needs, the specific ways a live session destroys them, a twelve-minute estimation round you can run this week, how to read the spread rather than just the average, and the categories of question where you should ignore the crowd entirely and go and ask one person who knows.

The four conditions the effect depends on

The classic framing sets out four requirements: diversity of opinion, independence, decentralisation, and a mechanism for aggregation. Drop any one and the crowd degrades toward being a single loud opinion with extra steps.

Diversity means people bring different information and different errors. A room of eight people from the same team who read the same dashboard yesterday is one estimate wearing eight faces. Independence means each judgement is formed without knowledge of the others. Decentralisation means people can draw on local, specific knowledge rather than a single shared briefing. Aggregation means there is an actual rule for combining the answers, decided before you see them.

In a workplace session, diversity and decentralisation are usually adequate by accident — people genuinely do know different things. Independence is the one you break, and aggregation is the one you forget to define, which lets whoever runs the meeting pick whichever summary suits the conclusion they already had.

Diversity of information

Invite people whose errors differ: someone from support, someone from delivery, someone who has done it before elsewhere. Eight people who share a manager share a bias.

Independence

Nobody sees or hears another estimate before submitting their own. This is the condition that silence and simultaneous reveal exist to protect.

Decentralised knowledge

Let people use what they specifically know rather than reasoning only from the brief you just presented. Ask for the estimate before you present your framing, not after.

A stated aggregation rule

Decide before the reveal whether you will use the median, the mean, or the middle 50 per cent, and say so. Choosing afterwards is how a group vote becomes a rubber stamp.

How a normal meeting destroys independence

The failure is fast and almost invisible. Someone asks 'roughly how long will this take?', the most senior or most confident person says 'six weeks', and every estimate after that clusters around six weeks. This is the anchoring effect operating on a number nobody had reason to trust, and it happens whether or not the anchor was intended as a suggestion.

Visible live results do the same thing more politely. If a poll shows running totals as votes arrive, later voters see a leader and drift toward it — the bandwagon effect — and by the end you have a majority that was partly manufactured by the display. The votes are real; the independence is not.

Seniority adds a third layer. Where a manager has stated a view, disagreement now has a cost, and what you get back is not an estimate but a negotiation. Groupthink research describes this as the suppression of dissent in favour of a comfortable consensus, and the resulting number can be worse than the worst individual guess, because everyone has abandoned their own information in favour of one person's.

The practical rule that follows: the person with the most authority estimates last, or does not estimate at all. If you are running the session and you have a number in your head, write it down privately and do not say it until the reveal.

A twelve-minute estimation round

1

Minute 0–2 — state the question precisely

Ambiguity in the question shows up as spread in the answers. 'How long will the migration take?' produces chaos; 'how many working days from kick-off to the first customer on the new system, assuming current staffing' produces comparable numbers. Write the exact wording on screen and leave it there.

2

Minute 2–3 — name the aggregation rule

Say out loud what you will do with the answers before you see them. For most workplace estimates: 'we will use the median, and we will look at the middle range rather than the extremes'. This one sentence removes the temptation to reinterpret the data once it is on screen.

3

Minute 3–6 — silent independent submission

Everyone submits a number with no discussion and no visible running results. Three minutes is enough for a considered guess and short enough that nobody starts researching. Anonymous submission is worth it here: it protects juniors from self-censoring toward the boss's likely view.

4

Minute 6–8 — reveal the whole distribution at once

Show every estimate together, not a summary. The shape is the interesting part: a tight cluster means shared understanding, a bimodal split means two groups are answering two different questions, and a long tail usually means one person knows something the rest do not.

5

Minute 8–11 — interrogate the outliers, not the average

Ask the highest and lowest estimator to explain their reasoning in sixty seconds each. Do not ask them to defend the number — ask what they were including. This is where the actual information transfer happens, and it is the only part of the exercise where discussion helps.

6

Minute 11–12 — decide whether to re-estimate

If the outliers surfaced a real constraint the group had not considered, run one more silent round. If they only surfaced different assumptions about scope, fix the scope wording instead and re-run. Never average round one and round two together.

8–12
Independent estimators
Enough for errors to cancel on most workplace questions; past about twenty the accuracy gain flattens while the coordination cost keeps rising.
3 min
Silent submission window
Long enough for a considered estimate, short enough that nobody starts building a spreadsheet mid-session.
2
Rounds maximum
One silent round, one optional re-estimate after the outliers explain. A third round is negotiation, not estimation.
60 sec
Per outlier explanation
Enough to say what they included that others did not, without turning into a defence of the number.

Median, mean, and reading the spread

For most estimates the median is the safer summary. Workplace guesses are usually skewed — nobody estimates below zero, but someone always estimates four times the median — and the mean gets dragged by that tail. The median ignores it, which is what you want when the tail is one person being pessimistic rather than one person being right.

The exception is when you suspect the tail carries information. If the one high estimator is the only person who has actually done this migration before, the median is confidently averaging away the best-informed answer in the room. That is why the outlier conversation comes before the decision, not after.

The spread deserves more attention than it usually gets. A tight cluster means the group shares a mental model, which is reassuring but also worth checking — tight agreement among people who all read the same document is not independent confirmation. A wide spread almost always means the question was ambiguous, and the fix is rewording rather than more discussion.

A bimodal distribution — two clumps with a gap in the middle — is the most actionable shape. It nearly always means two subgroups are estimating two different scopes. Name it out loud, ask each clump what they included, and you will usually find a genuine disagreement about the work that would otherwise have surfaced three weeks into delivery.

Open discussion first

  • Senior person names a figure in the first minute
  • Everyone else adjusts toward it
  • Spread looks reassuringly tight
  • Quiet dissent never enters the record
  • One opinion, eight signatures

Silent estimates first

  • Everyone commits before hearing anyone
  • Full distribution revealed at once
  • Outliers explain what they included
  • Scope ambiguity surfaces in minute eight
  • Aggregate rests on independent information

The two-round variant for harder questions

For high-stakes or genuinely uncertain questions, a compressed Delphi structure is worth the extra time. The Delphi method collects anonymous estimates, feeds back a summary of the group's responses and reasoning, and collects a second anonymous round — deliberately avoiding face-to-face debate so that authority and volume do not drive the convergence.

The workshop version fits in twenty-five minutes. Round one is silent and anonymous. You then share the distribution and, crucially, the anonymised reasoning — two or three sentences from the extremes and one from the middle — without naming who said what. Round two is again silent and anonymous, and people may move or hold.

Convergence in round two means the reasoning was genuinely persuasive. Persistent disagreement means you have a real unknown, and the correct output of the session is not a number but a named question to go and answer. Recording that outcome honestly is more valuable than manufacturing a figure everyone can live with, which is precisely how the Abilene paradox produces decisions nobody actually wanted.

A note on anonymity: it makes the round more honest but harder to follow up. Keep the estimates anonymous and the follow-up actions named, so the candour survives without the accountability disappearing with it.

When the crowd is worse than one expert

Aggregation helps when everyone has partial information about the same thing. It does not manufacture knowledge nobody has. For questions where the answer depends on specific technical facts — whether a particular database will hold under a given load, what a regulation actually requires — twelve confident guesses average to a confident wrong answer, and the confidence is the dangerous part.

The test is simple: could a well-informed person in this room plausibly have relevant partial information? If yes, aggregate. If the honest answer is that only one or two people could possibly know, do not run the poll — asking a room to vote on a fact is theatre, and it borrows the credibility of the group to launder a guess.

Crowds also degrade when the group has been briefed into a single view. If everyone just sat through the same twenty-minute presentation making the case for six weeks, their estimates are not independent regardless of how you collect them. Where possible, collect the estimate before the presentation, or accept that you are measuring how persuasive the presentation was.

Finally, be careful about aggregating preferences and calling it accuracy. A vote on which option the team prefers is a legitimate decision mechanism. It is not the wisdom of the crowd, because there is no true value the errors are scattering around, and treating a preference vote as evidence of correctness is a common and expensive confusion.

Practical setups that hold the conditions

Top Tips

  • Turn off live results for any estimation question. Session Flo can hold responses and reveal the full distribution at once, which is the single change that does most to protect independence.
  • Ask for the number before you present your framing. Once you have made the case, you are measuring persuasion rather than judgement, and you cannot get the original independence back in the same session.
  • Put the exact question wording on screen and leave it there for the whole submission window. Most wide spreads are caused by three people quietly answering slightly different questions.
  • Have the most senior person submit last or abstain entirely, and never let them think aloud beforehand. Where that is politically impossible, run the round anonymously and say clearly that you will not be attributing estimates.
  • Use ranges rather than points for anything beyond a few weeks out. Asking for a best case and a worst case gives you the spread inside each head as well as across the room, and people are noticeably more honest about the worst case when it is a required field.
  • Log the estimate, the spread and the date somewhere durable. A team that can see its last six estimation rounds against what actually happened calibrates faster than any amount of discussion about estimating better.
"
A room that agrees within thirty seconds has not confirmed anything. It has heard one number and repeated it eleven times.

What to do with the result

An aggregated estimate is an input, not a verdict. The output of a good round is three things: a central figure, a spread, and a list of the assumptions that separated the extremes. Carrying only the first into the next meeting throws away most of what you produced.

Write the spread into whatever you commit to. 'Median twelve weeks, with the range running eight to twenty depending on whether the data migration is in scope' is a far more useful sentence than 'the team estimated twelve weeks', and it survives contact with a stakeholder who wants certainty you do not have.

Then close the loop. When the real answer arrives, show the group their original distribution next to it. Teams that see their own calibration improve at estimating in a way that no methodology change achieves on its own, and it takes about four minutes at the start of a session that was happening anyway.

Frequently asked questions

How many people do you need for the wisdom of the crowd to work?

Eight to twelve independent estimators is enough for most workplace questions. The gain from additional people flattens fairly quickly, and past about twenty the coordination cost outweighs the marginal accuracy. Independence and diversity matter far more than headcount: five people with genuinely different information beat thirty who all attended the same briefing.

Should I use the mean or the median of the estimates?

The median for most workplace estimates, because guesses are usually skewed and a single very high figure drags the mean without adding accuracy. Use the mean only when the distribution is roughly symmetric. Whichever you choose, state it before the reveal — picking the summary after you have seen the numbers turns the exercise into a way of justifying a conclusion you already held.

Does anonymity actually change the estimates?

It changes who is willing to submit an unpopular number. Anonymity removes the cost of disagreeing with a senior colleague, which is the main reason estimates cluster artificially. Keep the estimates anonymous and the follow-up actions named, so you get the candour without losing accountability for what happens next.

What does a very wide spread tell me?

Almost always that the question was ambiguous rather than that the group is bad at estimating. Before re-running, check whether people were including the same scope, the same start point and the same staffing assumption. A bimodal split — two clusters with a gap — nearly always means two subgroups estimated two different pieces of work.

When should I ignore the group and ask one expert?

When the answer depends on specific technical or regulatory facts that most of the room cannot plausibly hold partial information about. Aggregation cancels scattered errors; it cannot create knowledge nobody has. Asking twelve people to vote on a fact produces a confident average that is confidently wrong, and the group's apparent agreement makes it harder to challenge.

Is a team vote on options the same as the wisdom of the crowd?

No. A vote on which option people prefer is a decision mechanism, and a perfectly good one. The wisdom of the crowd applies to questions with a true value that individual errors scatter around. Treating a preference vote as evidence that the chosen option is correct confuses popularity with accuracy, and it is a common and expensive mistake.

Keep reading

Related articles

Ready to run better sessions?

Create interactive events with live polls, quizzes, icebreakers, and more.