Data Sources

FIRE Lab's simulations are only as good as the history behind them. This page documents every dataset the simulator uses, how recent years are extended, and the known limitations you should keep in mind when reading results.

JST Macrohistory Database (16 countries, 1871–2020)

The primary dataset is the Jordà-Schularick-Taylor Macrohistory Database (release 6), an academic panel covering 18 advanced economies with annual data on equity total returns, government bond returns, short-term bills, consumer price inflation, house prices, and exchange rates. FIRE Lab uses 16 countries with sufficiently complete return histories, starting in 1871. The market-history presets default to 1900 because that is the first year all 16 countries are covered (Switzerland, Spain, and the Netherlands begin in 1900).

JST is the standard academic source for long-run cross-country return data, underpinning the well-known "Rate of Return on Everything" research. All returns are converted to real (inflation-adjusted) local-currency terms before simulation.

The database is published under a CC BY-NC-SA 4.0 license by the original authors (Jordà, Schularick, Taylor et al.).

2021–2025 extension (unofficial)

Official JST data ends in 2020. FIRE Lab extends each country through 2025 using public sources: equity returns from national total-return indices via Yahoo Finance (priced on annual averages to match JST's convention), CPI from the IMF World Economic Outlook, and bond returns reconstructed from OECD long-term interest rates.

The extension follows JST's methodology as closely as practical — including annual-average equity pricing and legacy-currency handling for Eurozone members — and is validated against overlapping official data. It is nonetheless unofficial: treat the most recent five years as best-effort estimates rather than curated academic data.

US dataset (Bogleheads / Simba, 1871–2025)

The US-only mode uses long-run US market data as organized by the Bogleheads community's Simba backtesting spreadsheet: S&P 500 total returns with pre-1957 predecessor indices (ultimately Shiller's 1871+ series), MSCI EAFE for international equities from 1970, and 10-year US Treasury returns.

This dataset reflects the single most successful major equity market of the past 150 years. Results on US-only data are therefore an optimistic benchmark; the global pooled mode is the more conservative planning baseline.

Pre-1970 international backfill

The US dataset variant with international backfill replaces the pre-1970 international column (which is a US placeholder in the original spreadsheet) with a GDP-weighted ex-US equity series derived from JST, restoring genuine US/non-US diversification in early history. Per-country JST returns have been validated against corresponding MSCI indices (differences within ~0.3pp annualized).

Processing and sampling

All simulations run on real returns. Multi-country simulations pool the 16 countries with equal probability (1/N) per sampled block — no GDP weighting — a choice validated by sensitivity analysis showing safe-withdrawal-rate differences under 0.1pp across weighting schemes.

Block Bootstrap sampling draws contiguous blocks of 5–15 years (uniform) with same-country circular wrap-around, preserving multi-year momentum and volatility clustering. See the methodology page for details.

Known limitations

Survivorship and selection: the 16 countries are economies that remained investable for 150 years. Markets that closed permanently (Russia 1917, China 1949) are not represented, which biases all long-run datasets optimistic.

House-price indices are appraisal-based and smoothed, understating true volatility; they inform the buy-vs-rent tool but are not used as an investable asset in portfolio simulations.

Historical returns do not guarantee future performance. Taxes, currency risk for international investors, and trading frictions beyond fund expense ratios are not modeled.


Back to the simulator · Methodology · FAQ