A backtest is evidence, not proof.
Historical testing can help determine whether a systematic market hypothesis deserves further investigation, but a strong historical result does not establish that the same behavior will persist. AEC treats backtests as research instruments whose usefulness depends on the quality of the hypothesis, data, implementation and validation process.
Separate development from evaluation.
In-sample observations can support hypothesis development and parameter selection. Where the research design permits, separate out-of-sample or forward-style evaluation provides a more demanding test of whether observed behavior survives outside the environment used to construct the strategy.
Repeatedly modifying a model after observing supposedly out-of-sample results can gradually turn that sample into another development set. Research records should therefore preserve when decisions were made and which information was available at the time.
Look for neighborhoods, not isolated peaks.
A configuration that performs well only at one narrow parameter value may reflect noise rather than a durable market relationship. Robustness analysis can compare neighboring settings and examine whether economically similar configurations produce reasonably consistent behavior.
The objective is not to find parameters that make every historical period profitable. It is to determine whether the underlying hypothesis remains recognizable when reasonable implementation choices change.
Challenge the strategy across periods and regimes.
Aggregate statistics can conceal dependence on a particular market environment. AEC research can segment results by period, volatility, trend behavior, market structure, session and other relevant conditions to determine where a strategy historically succeeded or failed.
Regime analysis is descriptive rather than predictive. A strategy that behaved consistently across several historical environments can still deteriorate when market structure changes.
Stress costs, fills and implementation assumptions.
Gross historical performance can overstate what an executable strategy might have achieved. Commissions, slippage, contract sizing, fill assumptions, latency and liquidity should be incorporated at levels appropriate to the strategy being evaluated.
Researchers can also increase assumed friction beyond a base case to determine how quickly the historical edge deteriorates. A thesis that disappears under modest execution stress deserves additional scrutiny.
Study losing periods instead of hiding them.
Drawdowns, losing clusters and failed signals are part of the evidence. Reviewing them can help distinguish ordinary variance from weaknesses in the hypothesis, implementation or risk framework.
Negative tests should remain part of the research record. Preserving failed experiments reduces the risk of repeatedly rediscovering rejected ideas and provides context for later modifications.
Promotion depends on the body of evidence.
No single statistic establishes robustness. Profit factor, win rate, net results, drawdown and trade count each describe different characteristics and should be considered alongside stability, costs, regime behavior and the economic rationale for the strategy.
AEC distinguishes exploratory research from configurations that have survived stronger validation. Historical backtests and simulated results remain separately classified from any actual executed trading record.