A bee at the golden centre of a purple water lily

Every strategy is closely examined by the QSL Strategy Validation Process.

Strategies need to be tested in multiple ways to give us a better understanding of their strengths and weaknesses. Here is a brief overview of the validation process we will teach in our free course ‘Evidence-Based Investing for Everyone’.

01The Problem To Be Aware Of

Backtests can be manufactured to look good. That is called overfitting.

A backtest has charts and tables showing how a strategy would have performed on historical data. That is a great tool. But when we have a backtest that looks good, there could be a common reason for that: overfitting. The strategy rules have been adjusted many times until the chart looks good.

A strategy can fail with real money for reasons that were present in the backtest all along: too little history to measure anything, trading costs left out, or settings adjusted until that one stretch of data looked good. Each of those has a name, and a test that catches it.

This strategy delivered a fantastic backtest. The annual return was over 60% in the period from Jan 2020 to Jan 2024. It outperformed the S&P 500 (blue line). But can we trust this backtest? The QSL Strategy Validation Process will help us to answer this question.

Exhibit — a backtest built to look good

vs S&P 500 buy-and-hold

Backtest equity curve growing from $100,000 to about $680,000 between January 2020 and January 2024, far above the S&P 500 benchmark line, with monthly returns below
Hypothetical backtested performance, shown as a teaching example. See the disclosures at the foot of this page.
02The Core Idea

The QSL Strategy Validation Process

Before anyone should trust a strategy, four questions have to be answered, and they have to be answered in that order. Each one only means something once the question before it has been settled, so a result is reported as the point a strategy reached rather than as a single pass or fail. That is more useful, because it tells you what to work on next.

These tests look for what professionals call an edge— a strategy’s real, repeatable advantage over holding the market. Whether an edge exists, and whether it is good enough to matter, is what the middle of the process measures.

The four questions break down into thirteen individual tests, grouped into the four levels below. Open a level to see its tests — and, on the right of each, how the backtest from the previous section actually did.

Level 13 tests

Validity of the test

Is this result measurable at all?

Clear this level and you know: the test was set up honestly, so its numbers mean something.

Show the 3 tests
01

What must this beat, and what is it for?The professionals call this the Benchmark & Objective Declaration.

Passed

Before any testing, we wrote down what the strategy had to beat — the S&P 500 — and what it was for. Deciding that up front, rather than afterwards, keeps every test that follows honest.

02

Can we trust our data?The textbook name is the Data Integrity test.

Passed

The price history is complete and the code cannot peek at tomorrow’s prices. The numbers underneath everything else can be trusted.

03

Do we have enough transactions?In statistics this is called Sample Adequacy.

Passed

Four years and hundreds of trades is comfortably enough history for the strategy’s record to mean something rather than being a fluke of a handful of trades.

Level 23 tests

Existence of the edge

Is there a real edge here?

Clear this level and you know: there is something real to evaluate: the result does not disappear once trading costs are counted, and it is not luck.

Show the 3 tests
04

Does it make money on average?Practitioners call this Expectancy.

Passed

On a typical trade the strategy made money, and there is enough history to be confident that average is real rather than luck.

05

Does it survive the cost of trading?In a research report this appears as Transaction-Cost Survival.

Passed

The profit is far larger than the cost of buying and selling — it still holds up even when we charge several times the realistic cost.

06

Could random trading have done this?The technical name is Statistical Significance.

Passed

On the data it was built from, a result this good would almost never happen by chance. (On fresh data later it was no better than a coin toss — an early warning of what comes below.)

Level 34 tests

Quality and durability of the edge

Is the edge good, and is it the strategy’s own?

Clear this level and you know: the edge lasts, survives bad conditions, is worth its risk, and is not merely the result of share prices rising in general.

Show the 4 tests
07

Did it make money in every period?Analysts know this as Temporal Stability.

Failed

The edge was strong in the early years and had faded to nothing by the recent ones. Something that only worked in the past is not worth trusting now.

08

Does it survive crashes and bear markets?In research papers this is called Regime Robustness.

Caution

In the worst market stretches the strategy fell much harder than a “similar risk” strategy should — close to half its value at the low point.

09

Is it worth the risk?The industry term is Risk-Adjusted Performance.

Caution

It did beat the market on the old data, but only by riding much bigger ups and downs to get there.

10

Does the return come from the strategy rule, or from the fact that all stocks moved up in the backtest period?Quantitative researchers call this Factor Attribution.

Failed

Most of the gains turned out to be ordinary market movement in disguise, not a skill the strategy itself added.

Level 43 tests

Overfitting controls

Was the result manufactured by the research process?

Clear this level and you know: the result reflects how prices actually behaved, not the settings the researcher chose or the number of attempts they made.

Show the 3 tests
11

Does it still work if we change the settings?Researchers refer to this as Parameter Sensitivity.

Passed

Nudging the strategy’s settings up and down barely changed the result — a sign it found a broad sweet spot rather than one lucky setting.

12

Does it keep working on new data?In backtesting this is known as Out-of-Sample Validation.

Failed

On data it had never seen, the strategy lost money while the market rose. This is the decisive test, and it failed.

13

How many tries were behind this result?The formal name is the Multiple-Testing Correction.

Unknowable

We were never told how many versions were tried before this one, so we cannot rule out that it was the best of many attempts. That leaves it unproven rather than passed.

03What We Cannot Promise

Honesty over comfort.

Validation means the backtest was constructed and evaluated honestly. It does not mean the future will cooperate.

A strategy can clear every test and still lose money.

Past performance is not a guarantee of future results.

Backtests are not the same as live trading.

Our validation is a starting point, not a finish line.

We cannot guarantee performance. But we can guarantee that every strategy you see on QS Lab earned its place.

04How We Work

What we commit to.

Show the working

Every strategy arrives with its validation record, so you can read how the result was reached rather than take it on trust.

Match the test to the strategy

A trend strategy and an income strategy fail in different ways, so they are not measured by an identical checklist.

Publish what failed

Strategies that do not clear the levels are published with the reason. A failure tells you something a success cannot.

Score, and explain the score

A result is a number out of 100 with the record behind it, not a bare pass or fail.

Keep watching after publication

Live data accumulates, and the record is updated as it does.

Ready to see what passes?

Get early access

Free. Self-paced.

Important Disclosures

All backtested and simulated results shown on this site are hypothetical performance results — they were achieved with the benefit of hindsight and do not represent actual trading or investment returns. Hypothetical performance results have many inherent limitations: they do not account for actual trading costs, bid-ask spreads, slippage, liquidity constraints, margin requirements, or taxes, and the results may not reflect conditions that would have prevailed during the periods shown. No representation is made that any individual trader will achieve results similar to those shown.

Past performance, whether actual or hypothetical, is not indicative of future results. All content on this platform is provided for informational and educational purposes only. Nothing here constitutes investment advice, a solicitation, or a recommendation to buy or sell any security or financial instrument. Independent traders should consult a qualified, licensed financial professional before making any investment decision.