
Every strategy is closely examined by the QSL Strategy Validation Process.
Strategies need to be tested in multiple ways to give us a better understanding of their strengths and weaknesses. Here is a brief overview of the validation process we will teach in our free course ‘Evidence-Based Investing for Everyone’.
Backtests can be manufactured to look good. That is called overfitting.
A backtest has charts and tables showing how a strategy would have performed on historical data. That is a great tool. But when we have a backtest that looks good, there could be a common reason for that: overfitting. The strategy rules have been adjusted many times until the chart looks good.
A strategy can fail with real money for reasons that were present in the backtest all along: too little history to measure anything, trading costs left out, or settings adjusted until that one stretch of data looked good. Each of those has a name, and a test that catches it.
This strategy delivered a fantastic backtest. The annual return was over 60% in the period from Jan 2020 to Jan 2024. It outperformed the S&P 500 (blue line). But can we trust this backtest? The QSL Strategy Validation Process will help us to answer this question.
Exhibit — a backtest built to look good
vs S&P 500 buy-and-hold

The QSL Strategy Validation Process
Before anyone should trust a strategy, four questions have to be answered, and they have to be answered in that order. Each one only means something once the question before it has been settled, so a result is reported as the point a strategy reached rather than as a single pass or fail. That is more useful, because it tells you what to work on next.
These tests look for what professionals call an edge— a strategy’s real, repeatable advantage over holding the market. Whether an edge exists, and whether it is good enough to matter, is what the middle of the process measures.
The four questions break down into thirteen individual tests, grouped into the four levels below. Open a level to see its tests — and, on the right of each, how the backtest from the previous section actually did.
Level 13 testsValidity of the test
Is this result measurable at all?
Clear this level and you know: the test was set up honestly, so its numbers mean something.
Show the 3 testsHide the tests
Validity of the test
Is this result measurable at all?
Clear this level and you know: the test was set up honestly, so its numbers mean something.
Show the 3 testsHide the testsWhat must this beat, and what is it for?The professionals call this the Benchmark & Objective Declaration.
Before any testing, we wrote down what the strategy had to beat — the S&P 500 — and what it was for. Deciding that up front, rather than afterwards, keeps every test that follows honest.
Can we trust our data?The textbook name is the Data Integrity test.
The price history is complete and the code cannot peek at tomorrow’s prices. The numbers underneath everything else can be trusted.
Do we have enough transactions?In statistics this is called Sample Adequacy.
Four years and hundreds of trades is comfortably enough history for the strategy’s record to mean something rather than being a fluke of a handful of trades.
Level 23 testsExistence of the edge
Is there a real edge here?
Clear this level and you know: there is something real to evaluate: the result does not disappear once trading costs are counted, and it is not luck.
Show the 3 testsHide the tests
Existence of the edge
Is there a real edge here?
Clear this level and you know: there is something real to evaluate: the result does not disappear once trading costs are counted, and it is not luck.
Show the 3 testsHide the testsDoes it make money on average?Practitioners call this Expectancy.
On a typical trade the strategy made money, and there is enough history to be confident that average is real rather than luck.
Does it survive the cost of trading?In a research report this appears as Transaction-Cost Survival.
The profit is far larger than the cost of buying and selling — it still holds up even when we charge several times the realistic cost.
Could random trading have done this?The technical name is Statistical Significance.
On the data it was built from, a result this good would almost never happen by chance. (On fresh data later it was no better than a coin toss — an early warning of what comes below.)
Level 34 testsQuality and durability of the edge
Is the edge good, and is it the strategy’s own?
Clear this level and you know: the edge lasts, survives bad conditions, is worth its risk, and is not merely the result of share prices rising in general.
Show the 4 testsHide the tests
Quality and durability of the edge
Is the edge good, and is it the strategy’s own?
Clear this level and you know: the edge lasts, survives bad conditions, is worth its risk, and is not merely the result of share prices rising in general.
Show the 4 testsHide the testsDid it make money in every period?Analysts know this as Temporal Stability.
The edge was strong in the early years and had faded to nothing by the recent ones. Something that only worked in the past is not worth trusting now.
Does it survive crashes and bear markets?In research papers this is called Regime Robustness.
In the worst market stretches the strategy fell much harder than a “similar risk” strategy should — close to half its value at the low point.
Is it worth the risk?The industry term is Risk-Adjusted Performance.
It did beat the market on the old data, but only by riding much bigger ups and downs to get there.
Does the return come from the strategy rule, or from the fact that all stocks moved up in the backtest period?Quantitative researchers call this Factor Attribution.
Most of the gains turned out to be ordinary market movement in disguise, not a skill the strategy itself added.
Level 43 testsOverfitting controls
Was the result manufactured by the research process?
Clear this level and you know: the result reflects how prices actually behaved, not the settings the researcher chose or the number of attempts they made.
Show the 3 testsHide the tests
Overfitting controls
Was the result manufactured by the research process?
Clear this level and you know: the result reflects how prices actually behaved, not the settings the researcher chose or the number of attempts they made.
Show the 3 testsHide the testsDoes it still work if we change the settings?Researchers refer to this as Parameter Sensitivity.
Nudging the strategy’s settings up and down barely changed the result — a sign it found a broad sweet spot rather than one lucky setting.
Does it keep working on new data?In backtesting this is known as Out-of-Sample Validation.
On data it had never seen, the strategy lost money while the market rose. This is the decisive test, and it failed.
How many tries were behind this result?The formal name is the Multiple-Testing Correction.
We were never told how many versions were tried before this one, so we cannot rule out that it was the best of many attempts. That leaves it unproven rather than passed.
Honesty over comfort.
Validation means the backtest was constructed and evaluated honestly. It does not mean the future will cooperate.
A strategy can clear every test and still lose money.
Past performance is not a guarantee of future results.
Backtests are not the same as live trading.
Our validation is a starting point, not a finish line.
We cannot guarantee performance. But we can guarantee that every strategy you see on QS Lab earned its place.
What we commit to.
Show the working
Every strategy arrives with its validation record, so you can read how the result was reached rather than take it on trust.
Match the test to the strategy
A trend strategy and an income strategy fail in different ways, so they are not measured by an identical checklist.
Publish what failed
Strategies that do not clear the levels are published with the reason. A failure tells you something a success cannot.
Score, and explain the score
A result is a number out of 100 with the record behind it, not a bare pass or fail.
Keep watching after publication
Live data accumulates, and the record is updated as it does.
Ready to see what passes?
Free. Self-paced.