Principles

Methodology

The rules that make the record worth reading. They describe how forecasts behave here — not the proprietary machinery behind any individual number.

Why probabilistic forecasts

A forecast is only meaningful if it can be wrong. Every claim here is expressed as a probability attached to a clearly resolvable question, so that — once reality decides — the claim can be scored without argument.

Forecasts versus views

A forecast is a quantitative, falsifiable, resolvable prediction (e.g. "≤ 3.50% on a stated date: 64%"). A view is qualitative interpretation. Only forecasts enter the scoring statistics. Commentary never inflates the track record.

Forecasts are distributions

A forecast is a probability distribution over outcomes, not a single number. Binary questions carry one probability; multiple-choice questions a probability over mutually exclusive options; numeric and date questions a full distribution over a declared range, from which thresholds like P(> 100) are derived rather than asked separately. Each type is scored on the representation it uses.

Resolution criteria

Each forecast names, in advance, the exact condition that will count as a "yes" and the source that will settle it. Resolution is mechanical, not a matter of later interpretation.

Timestamps

Forecasts carry ISO-8601 publication times, and enter the public Git history when committed. That history is an audit aid — it makes silent backdating visible — but it is not claimed to be an absolute tamper-proof mechanism. Where stronger proof is warranted, forecasts can be sealed with a cryptographic commitment.

Updates never overwrite

A forecast may be updated as evidence arrives. Updates add states; they do not erase the original. The current probability is simply the latest state. A forecast that aged badly stays on the record — that is what makes the record honest.

Scoring

Every type is scored with a proper scoring rule — one that rewards reporting your true probabilities rather than gaming the metric. The primary score is the logarithmic score, ln P(outcome), which applies across binary, multiple-choice, numeric and date forecasts; higher (less negative) is better, and it punishes confident mistakes hard. Alongside it: the Brier score for binary and multiple-choice questions, and CRPS for numeric distributions — both lower-is-better diagnostics. Because raw scores depend on how hard the questions were, each forecast is also reported against an uninformative baseline, which shows whether it carried real information. Everything is computed from the raw record, never entered by hand.

Calibration

Beyond average score, forecasts are grouped into probability buckets and compared against realized frequencies. Calibration is only meaningful with enough observations, so the number of resolved forecasts per bucket is always shown alongside it.

Benchmarks

Where a reasonable baseline exists — a naïve base rate, a market-implied probability, a community forecast — the record leaves room to compare against it. Benchmarks are recorded only when genuine; none are fabricated.

Question selection & cherry-picking

Scoring is only half the problem. A proper scoring rule rewards honest probabilities conditional on a question. It cannot stop a forecaster from choosing only easy questions and posting impressive averages that show little real skill.

So the record separates two things. A Core Track Record is generated from predeclared recurring cohorts: a fixed question universe, a fixed cadence, mandatory coverage, and weights set in advance. Every eligible question is forecast; a miss is recorded as a miss, not quietly dropped. Discretionary Featured forecasts stay public and fully scored, but are reported separately and never pooled into the Core record. Threshold probabilities read off a single distribution — for example P(Brent > 100) — are shown as derived views, not counted as independent forecasts.

Nothing is removed for being wrong. Demonstration data is clearly marked and excluded from every metric. The score evaluates the forecast; the selection protocol protects the record.

Relationship to any private system

This site is a public ledger. It may receive forecasts produced with the help of a separate, private forecasting process, but it stays decoupled from that process and exposes none of its internals. What is published here is exactly what is needed to judge the forecasts: the claims, the probabilities, the timing and the outcomes.