← The Evidence
Rule design7 min readJul 4, 2026

Why backtests lie, and how to read one honestly

Any strategy can be made to look brilliant on historical data. That isn’t a fact about the strategy, it’s a fact about backtests. Understanding why is the difference between a backtest that protects you and one that sets you up.

Here’s the uncomfortable part: a backtest can’t tell the difference between a real edge and a lucky fit. It just reports what would have happened. If you try enough variations on the same slice of history, one of them will look amazing, by chance alone, and the backtest will hand you that curve with a straight face.

The overfitting trap

This isn’t a vague worry; it’s a proven result. Bailey, Borwein, López de Prado, and Zhu showed that with enough trials you are mathematically guaranteedto find a strategy with an impressive backtested Sharpe ratio and zero genuine edge. They call it backtest overfitting, and they derive a “minimum backtest length”: the more configurations you try, the longer a track record you need before a good result means anything at all.

Every time you nudge a threshold to make the equity curve smoother, 15% instead of 20%, a 40-day average instead of 50, you are, quietly, fitting the noise in one particular history. The curve gets prettier. The edge gets more imaginary.

The history has already been strip-mined

It gets worse: you’re not the first to try. Harvey, Liu, and Zhu reviewed the hundreds of “factors” published in the finance literature and argued that, once you account for how many strategies have been tested against the same market history, most published edges are probably false positives.

t > 3.0
The statistical bar they argue a new strategy should clear, well above the textbook 2.0, precisely because thousands of researchers have data-mined the same past. Most “edges” don’t come close. (Harvey, Liu & Zhu, 2016)

There is exactly one history of the market, and it has been searched exhaustively. A backtest that looks great is, more often than not, a coincidence someone was always going to stumble onto.

Why the pretty curve breaks live

The flexibility that let a rule hug twenty years of data is the same flexibility that shatters when the future refuses to rhyme with the past. An over-optimized backtest is a detailed, confident description of a world that no longer exists. It tells you about yesterday with false precision and about tomorrow not at all.

What an honest backtest does instead

None of this means throw the backtest away. It means change what you ask of it. An honest backtest doesn’t sell you a curve, it stress-tests a rule you already believe in. It reports what the rule would have done in plain terms (how often it fired, what it would have triggered, the effect on returns), and it actively flags the ways it could be fooling you: thresholds tuned too tightly to the past, a rule that only ever works in one market regime, positions that look diversified but are the same bet.

The honest version has fewer knobs, resists the urge to curve-fit, and treats a suspiciously beautiful result as a warning rather than a selling point. That’s the line between a backtest as a sales tool and a backtest as a reality check.

It’s why Cetagon backtests every rule, to talk you out of the fragile ones, not into them, and why an AI critique reviews each rule for exactly the overfitting and regime-sensitivity this essay is about. Pressure-test a rule →

Sources

  1. Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the American Mathematical Society.
  2. Harvey, C. R., Liu, Y., & Zhu, H. (2016). …and the Cross-Section of Expected Returns. Review of Financial Studies.
  3. López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley. (Backtest overfitting; the deflated Sharpe ratio.)