The Good Judgment Project ran from 2011 to 2015 as part of an IARPA forecasting tournament. Thousands of volunteers answered nearly five hundred questions about world events and produced well over a million individual judgments, each one a probability attached to a precisely defined outcome.
The design is the interesting part. Questions had to be written so that a resolution could not be argued about later. Every forecast was scored with the Brier score. Nobody was graded on how confident they sounded or how good the reasoning looked in hindsight.
The results held up under scrutiny. Trained volunteers, with no classified information, outperformed comparison groups by wide margins, and the top performers kept performing: about 70% of superforecasters retained the status year over year, with a 0.65 correlation in performance between consecutive years. When the organizers repeated the exercise with later cohorts, the effect reappeared.
Three design rules transfer almost unchanged to a university challenge. Define the question before anyone answers it. Score with a rule that punishes overconfidence as well as error. Publish the scores, so participants can see calibration improve or fail to improve over time.
A leaderboard built on returns alone rewards whoever took the most risk in a lucky month. A leaderboard that also scores stated probabilities against outcomes measures something closer to skill, and it gives a professor a defensible basis for a grade.
Educational material. Not investment advice.