Case StudyMarch 6, 2026·9 min read

Case Study: A Forecasting Tournament With a Million Judgments

How the Good Judgment Project scored thousands of volunteers, and what its rules suggest for scoring a student competition.

The Good Judgment Project ran from 2011 to 2015 as part of an IARPA forecasting tournament. Thousands of volunteers answered nearly five hundred questions about world events and produced well over a million individual judgments, each one a probability attached to a precisely defined outcome.

The design is the interesting part. Questions had to be written so that a resolution could not be argued about later. Every forecast was scored with the Brier score. Nobody was graded on how confident they sounded or how good the reasoning looked in hindsight.

The results held up under scrutiny. Trained volunteers, with no classified information, outperformed comparison groups by wide margins, and the top performers kept performing: about 70% of superforecasters retained the status year over year, with a 0.65 correlation in performance between consecutive years. When the organizers repeated the exercise with later cohorts, the effect reappeared.

Three design rules transfer almost unchanged to a university challenge. Define the question before anyone answers it. Score with a rule that punishes overconfidence as well as error. Publish the scores, so participants can see calibration improve or fail to improve over time.

A leaderboard built on returns alone rewards whoever took the most risk in a lucky month. A leaderboard that also scores stated probabilities against outcomes measures something closer to skill, and it gives a professor a defensible basis for a grade.

Educational material. Not investment advice.

Sources

Every figure in this article comes from one of these publications. If something looks off, check the source before you take our word for it.

  1. Superforecasting: The Art and Science of PredictionPhilip E. Tetlock and Dan Gardner, Crown · 2015
  2. Evidence on good forecasting practices from the Good Judgment ProjectAI Impacts · 2019
  3. Verification of Forecasts Expressed in Terms of ProbabilityMonthly Weather Review 78(1), Glenn W. Brier · 1950

In the meantime

The platform is ready. A 30-minute demo shows more than any article.

Related resources

Case StudyCase Study: 66,465 Households at a Discount Broker

Barber and Odean matched trading records to returns from 1991 to 1996. The most active traders earned 11.4% while the market returned 17.9%.

8 min read
Case StudyCase Study: 91 Pension Plans and the 93.6% Myth

The most quoted number in asset allocation is also the most misquoted. What Brinson, Hood and Beebower actually measured, and what they did not.

10 min read
ATLAS – See the Odds Before You Invest