Why the testing lesson matters
The concept lesson explains the logic. This lesson asks whether that logic survives contact with actual rules, costs and failure conditions. The first discipline is to freeze the strategy before looking at the result. If entry, exit, sizing or exceptions are edited after a loss, the test is no longer evaluating the same strategy.
A useful strategy lab does not ask only “did it make money?” It asks whether the rule was executable, whether the costs were modeled honestly, whether the sample represented more than one favorable market regime, and whether the losses came from normal variance or from an assumption that no longer held.
Freeze the rule set
Write the strategy in a table before testing:
| Rule component | What must be fixed before the test |
|---|---|
| Market / universe | Which assets and venues are eligible, and why |
| Timeframe | Data interval and decision time |
| Entry | Exact observable condition that creates a position |
| Position size | How risk or allocation is calculated |
| Exit / invalidation | What closes or reduces the position |
| Costs | Fees, spread, slippage and other relevant friction |
| No-trade condition | When the signal is ignored even if it appears |
| Review trigger | Evidence that causes investigation rather than ad-hoc editing |
For expectancy, the most important assumptions from the paired lesson should appear explicitly in this table. If an assumption cannot be observed or tested, label it as judgment rather than pretending it is mechanical.
Build the cost and friction ledger
Start with gross strategy outcome and then subtract the costs created by the way the strategy actually trades. A simple educational ledger is:
Net result = Gross trading result − explicit fees − estimated spread/slippage − financing or transfer costs − other strategy-specific friction.
Not every family has every cost. Long-term investing may have low turnover but meaningful custody and conversion considerations. Arbitrage may be dominated by execution and transfer costs. Market making may depend on queue position, adverse selection and inventory. The purpose is to model the costs that belong to this strategy rather than paste one generic fee assumption across all styles.
Stress the assumption that matters most

Expectancy can be distorted by one huge outlier, a small sample, regime concentration or underestimated costs. A positive historical value is not a guarantee of future profit. Confidence should increase only when the strategy continues to produce similar behavior across new data and relevant market regimes.
The trader should also monitor realized R, not just planned R. If losses routinely exceed -1R or winners are cut early, the live expectancy is different from the strategy on paper.
Turn that into at least three stress cases: normal, unfavorable but plausible, and assumption failure. The third case is especially important. A strategy should not survive every scenario by definition. If no observation can make the method invalid, the rule is belief rather than a testable strategy.
Run the family-specific strategy lab
Build a table of at least thirty fictional or historical rule-based trades. Calculate win rate, average win, average loss, average cost and expectancy. Recalculate after removing the best trade and after increasing costs by 50%. Then split the sample by market regime. The skill is to understand what is producing the expectancy and how fragile that source may be.
After the first pass, change one assumption at a time. Increase costs, worsen entry quality, remove the best trade, or test a different regime. Do not optimize until the original result disappears; the aim is to understand sensitivity. A robust idea should usually make sense across a reasonable neighborhood of assumptions, even if the exact result changes.
Separate strategy failure from trader failure
When a test disappoints, classify the problem before changing anything:
- Rule failure: the strategy did exactly what it was designed to do, but the edge was insufficient.
- Execution failure: the signal had value but realistic costs, delay or liquidity destroyed it.
- Regime mismatch: the strategy was used in conditions it was not designed for.
- Process failure: the tester changed rules, skipped trades or used information unavailable at the time.
- Insufficient evidence: the sample is too small or too concentrated to support a conclusion.
Those categories lead to different next steps. Treating all losses as “bad strategy” prevents learning; treating all losses as “bad luck” prevents accountability.
Define the decision before the next sample
End the worksheet with one of four states: continue testing, investigate, modify as a new version, or retire / do not use. If you modify a material rule, give the strategy a new version and restart the relevant evidence trail. Do not blend the new rules into the old backtest as if they had always existed.
Skill check — no money needed

You can complete this lesson entirely with historical, fictional or paper data. The skill is demonstrated when another reader can reproduce the rule, recalculate the costs, see the same failure conditions and understand why the final decision was made. Profitability is not required for the exercise to be successful; discovering that a strategy does not survive realistic conditions is useful knowledge.
A losing period can be normal variance, while a profitable period can hide a broken process; retiring a strategy requires evidence about its mechanism, execution, and whether its edge still exists.
*Cryptocurrency and virtual asset transactions are highly volatile and irreversible, may result in significant losses, and do not guarantee returns; customers should trade only after understanding the risks involved.