Skip to content
Longterm Wiki has not been actively maintained since May 2026. Some data is still updated occasionally, but please don't rely on this site being current or accurate.
Longterm Wiki

ForecastBench: Dynamic LLM Forecasting Benchmark

web
forecastbench.org·forecastbench.org

ForecastBench is a dynamic benchmark measuring LLM forecasting accuracy against human baselines, relevant to AI safety as forecasting ability serves as a proxy for general intelligence and helps track AI capability progress toward and beyond human-level performance.

Metadata

Importance: 62/100tool pagetool

Summary

ForecastBench is a contamination-free benchmark that evaluates LLM forecasting accuracy against human comparison groups, including superforecasters. It maintains both a baseline leaderboard (no tools) and a tournament leaderboard (with scaffolding/tools), and projects when LLMs will reach superforecaster-level performance.

Key Points

  • •Dynamic, contamination-free benchmark preventing LLMs from training on benchmark questions, ensuring valid capability measurement.
  • •Compares LLM forecasting performance against human baselines including superforecasters as a proxy for general intelligence.
  • •Dual leaderboards: baseline (raw model performance) and tournament (with tool use, fine-tuning, ensembling).
  • •Tracks historical progress in LLM forecasting capabilities and projects date of LLM-superforecaster parity.
  • •Open to public submissions, enabling broad participation in capability evaluation.

Cited by 2 pages

PageTypeQuality
Forecasting Research Institute (FRI)Organization55.0
ForecastBenchProject53.0

1 FactBase fact citing this source

EntityPropertyValueAs Of
ForecastBenchFounded DateSep 2024—

Cached Content Preview

HTTP 200Fetched Oct 3, 20260 KB
Tournament leaderboard
Resource ID: kb-c808dd961e2e3c1d | Stable ID: sid_WNgUtLr8jP