We spent a week building a trading bot for prop-firm futures. Before writing a single strategy, we built the thing most traders never build: a lie detector for our own backtests.
Here's the problem. Test 100 random strategies on any price history and several will look brilliant โ not because they work, but because with enough tries, luck produces beautiful equity curves. Every "look at my backtest" screenshot you've seen on social media is, statistically, more likely to be this than not.
1. Walk-forward. Pick parameters using only past data, trade the next quarter blind, repeat. If your strategy only works when it's allowed to see the answers, this kills it.
2. Purged cross-validation. Slice history into folds, tune on some, score on others โ with a gap ("embargo") around every boundary so indicators can't leak information across. Reveals whether the edge lives everywhere or in one lucky year.
3. CPCV + the Probability of Backtest Overfitting. The heavy machinery, from Lรณpez de Prado's Advances in Financial Machine Learning. Run every combination of train/test blocks, and count: how often does the optimizer's in-sample favorite underperform the median strategy out-of-sample? That count is PBO. A real edge scores under 20%. Random noise scores ~50%. We fed it pure random-walk data as a control: PBO 47%, deflated Sharpe 0 โ the machine correctly calls noise on noise.
There's a fourth number, the deflated Sharpe ratio, which asks: given how many things you tried, what Sharpe would pure luck have produced โ and are you above it? It counts your attempts, which means it punishes you for the ideas you tried and threw away. Almost nobody computes it, because almost nobody wants the answer.
Over the next five posts we'll show you everything we tried against these gates โ including eight ideas that looked great and failed, several of which you are probably trading right now.