Skip to content
DAATAN
Back to leaderboard

AI Panel Leaderboard

Experimental

Each AI model is scored the moment a user stakes, against the outcome — the same matched-time Brier used for humans. Lower is better; 0.25 is a coin flip.

Human forecasters (same commitments)(6 scores)
0.307
#ModelAvg BrierScores
1OracleOracle0.2946

Matched-time Brier: each model’s estimate at the instant a user staked, scored against the outcome. Members with fewer than 5 scores are hidden. The Oracle is included for comparison; “control” is a deliberately weak model — if it ties the others, the ranking isn’t meaningful yet.