← Charlie Yan

Working paper · 2026

The Market Already Knew: A Pre-Registered Falsification of TimesFM-3 on the SPY Implied-Volatility Surface

A time-series foundation model does not improve on the market's own implied-volatility surface. Registered on the model's release day, before any arm ran.

TimesFM-3 shipped on 2026-08-31. The study was registered that same day, before any arm ran, and the registration is published verbatim — including all seven dated amendments. That ordering is the point: a falsification is only worth reading if the bar was set before the result was known.

The finding is in the title. Against the surface the market was already quoting, the model does not help.

What re-runs from the repo alone

Every statistical result — loss tables, Diebold–Mariano tests, Model Confidence Sets, Mincer–Zarnowitz recalibration, per-year robustness, and the wing follow-up — recomputes from the archived forecasts shipped in data/. No GPU, no model download, no vendor account. Evaluation is deterministic, so a re-run overwrites the shipped numbers with identical ones.

Regenerating the forecasts themselves needs a second environment and roughly 1–2 hours per configuration on an M1 GPU. The paper’s cost section documents both throughput pitfalls.

Basis and scope

Screen-basis forecast-loss comparisons on end-of-day quoted surfaces. No fills, no P&L, no performance claims. One constant is imported from proprietary data — a 0.173 vol-point 25Δ/30d round-trip spread used to scale the wing follow-up — and the basis it comes from is conservative against the candidate.

Unaffiliated with Google. TimesFM weights were used under their respective licenses for research only.

forecasting · foundation models · implied volatility · pre-registration