Backtest vs Forward Test: Judging a Market Strategy
Two honest ways to judge a prediction-market idea — and why the one that looks best on paper is usually the one that lied to you.
Published 2026-10-07 · Data as of 2026-10-07 · Market & data intelligence · Educational, not advice.
A backtest replays a strategy against past market prices; a forward test runs it on prices it has never seen. Backtests flatter you because you built the rule knowing the answer. Forward tests are slower and uglier but honest. Judge any prediction-market idea by what it does on data it could not have been fitted to.
Every prediction-market strategy sounds smart in hindsight. The honest question is not whether an idea would have worked on the prices you already have. It is whether it keeps working on prices it has never seen. Those are two different tests, and confusing them is the most common way a good-looking rule turns into a losing one.
A prediction market is an exchange where people contract against each other, not against a house. Every YES contract has a buyer and a seller; the exchange matches them, takes a fee, and never takes the other side itself. The price is just where the last buyer and seller agreed, and it moves continuously until the outcome is known. So when you test a strategy, you are really testing a set of decisions about when to take a position at a given price against a counterparty. The price history is your laboratory — and it is very easy to cheat in that lab without noticing.
What a backtest actually is
A backtest replays your rule against historical prices. You take a market that already resolved, feed in the prices as they moved, and ask: if I had opened a position whenever my rule fired, where would I have ended up? It is fast, it is cheap, and it feels like proof. You can run a hundred variations in an afternoon.
The problem is that you built the rule while already knowing how those markets resolved. That knowledge leaks in, usually without any bad intent. You pick a threshold because it happens to catch the three big moves in your sample. You add a filter because it removes the one ugly loss. Each tweak makes the line on the chart prettier, and each tweak is quietly fitting the rule to the past rather than to anything real. This is overfitting, and a backtest cannot detect it — the backtest is the thing being overfit.
Backtests also tend to assume you could have traded at the price you see. On thin markets, a position large enough to matter moves the price against you, and the counterparty you needed may not have existed at that moment. A clean historical price can describe a trade nobody could actually have filled.
What a forward test adds
A forward test runs the finished rule on prices it was never shown. You freeze the strategy — no more tweaking — and then you watch it work on new markets as they unfold, or on a slice of history you deliberately hid from yourself while building it. The point is simple: the data has to be information the rule could not have been fitted to.
Forward testing is slower and far less flattering. Returns shrink. Hit rates drift toward the middle. Edges that looked sharp on paper turn faint or vanish. That deflation is not failure — it is the measurement finally telling the truth. A rule that survives a forward test with most of its edge intact is worth attention. A rule that only shines in the backtest was describing the past, not predicting the future.
Walk-forward: the honest middle ground
You do not have to wait months to forward test. Walk-forward analysis splits history into blocks: you build the rule on an early window, test it on the next untouched window, then roll both forward and repeat. It approximates live conditions because every test block was genuinely out-of-sample when the rule met it. It is not a substitute for watching real prices move, but it catches the worst overfitting before you risk anything.
Reading your own results honestly
A few habits separate an honest evaluation from a flattering one. Count your trials — if you tried forty rule variants and reported the best, that winner is partly luck, and you should expect it to fade. Keep a true holdout you never touch until the very end. Account for fees and for the price you would actually have paid against a real counterparty, not the mid-price on the screen. And treat small samples with suspicion: a handful of resolved markets can make noise look like skill.
None of this tells you what to trade. It tells you how much to trust a number before you lean on it. The strategies that last are the ones whose edge survives being measured on data they could not have seen — and that is a test most ideas quietly fail.
Delta Arc's Prediction Markets product puts the live board from Kalshi and Polymarket in one place, which makes the raw material for this kind of honest evaluation — the same underlying outcome, priced by two different crowds — easy to watch as it moves. The next question worth asking: when two venues quote the same outcome at different prices, what is the gap actually telling you?
This is the free read. Delta Arc Prediction Markets shows you every top market across Kalshi and Polymarket in one live view. Get early access.