Back to leaderboard
AI Panel Leaderboard
ExperimentalEach AI model is scored the moment a user stakes, against the outcome — the same matched-time Brier used for humans. Lower is better; 0.25 is a coin flip.
Human forecasters (same commitments)(6 scores)
0.307| # | Model | Avg Brier | Scores |
|---|---|---|---|
| 1 | OracleOracle | 0.294 | 6 |
Matched-time Brier: each model’s estimate at the instant a user staked, scored against the outcome. Members with fewer than 5 scores are hidden. The Oracle is included for comparison; “control” is a deliberately weak model — if it ties the others, the ranking isn’t meaningful yet.