Test it on the week that hadn’t happened yet.
A single backtest is the easiest lie in this business to tell yourself. Fit a strategy to all of history, tune until the equity curve looks beautiful, and you have proven exactly one thing: that it can memorize the past. The lab exists to take the answers away.
Illustrative geometry, not results. An edge has to survive the window it was never fitted to — the ones that don’t get retired.
Optimize on a training window the strategy may learn from. Run it untouched on the window immediately after. Record what happened. Roll both forward, and repeat across years and regimes.
The bar is written down first.
A target set after the results are in is not a target. The lab states the goal and its thresholds before the search begins, then measures against them and reports what it found — including, and especially, when the answer is nothing.

- Net out-of-sample edge
- Positive after the round-turn cost of actually trading it. An edge that only survives frictionless trading is not one.
- Deflated Sharpe
- Adjusted for how many things were tried. Run enough variants and something will look good by accident; this is the correction for that.
- Probability of overfitting
- An estimate of how likely the selected configuration is to be a fit to noise rather than to structure.
- Out-of-sample evidence
- At least one full window the strategy was never allowed to learn from, replayed exactly.
Before a strategy is worth walking forward, a signal has to be worth testing. The last go/no-go screen put 143 candidates — 24 features across six horizons — through five gates: net edge above the cost hurdle, a sign that stays stable across halves, enough effective samples, a block-permutation p-value, and a multiple-comparison correction over all 143 trials. Nothing survived. The best deflated Sharpe was 0.63 against a 0.95 bar, and the count of near-misses sat inside the range the same machinery produces on pure noise.
The more useful finding was about the bar itself: the cost hurdle we had been quoting was commission only and ignored the spread, so everything had been scored against a floor that was too low. Neither of those is the kind of thing a lab that only reports its winners would ever find.
This makes cheating visible, not impossible.
You can still overfit the walk-forward process itself by running it enough times and keeping the run you liked. So the lab’s output is not a green light — it is evidence, and a person decides what that evidence is worth. The point is not to be certain. It is to be honest about how little a backtest proves.