Author: SERGII / Evidence-Based Trading Lab
Evidence status: DIAGNOSTIC — this note describes a research protocol, not a confirmed profitable strategy.
Imagine a corridor with 23,824 doors. Behind each one is a different threshold, lookback, asset filter, holding period, or entry rule. If we keep opening doors long enough, eventually one room will contain an impressive equity curve.
That curve may be a discovery. It may also be the inevitable winner of a large search.
The corridor is a metaphor; the number is not. In one branch of our research, the minimum recorded multiplicity budget reached 23,824 variants. They were not 23,824 independent experiments, so the number is not a complete statistical correction. It is a warning label: the wider the search, the easier it becomes for chance to resemble structure.
This is the first QuantConnect note from Evidence-Based Trading Lab. The project asks a narrow question: can a trading strategy earn the right to be called evidence by surviving rules that were locked before the result was known?
A backtest is an observation, not a verdict
The most dangerous strategy is not the obviously bad one. It is the strategy that looks good because the research process quietly helped it win.
We can change a threshold after seeing the result, choose a favorable interval, retain the best instrument from a large universe, assume a fill at the same close that created the signal, reduce costs, or inspect the final sample and continue calling it “untouched.” Each change may look defensible on its own. Together they turn an experiment into a story written backward.
Our defense is procedural. Before the decisive run, we freeze the question, data contract, execution model, cost model, comparison family, and failure conditions.
The seven gates
Hypothesis → Rule freeze → Development → Calibration → Locked final → Stress tests → Verdict
1. Hypothesis. State what should happen and what result would count as failure.
2. Rule freeze. Lock the data identity, signal timing, order timing, costs, metrics, and comparison set.
3. Development. Make the idea operational. This stage is allowed to influence the design.
4. Calibration. Check whether the result depends on one convenient historical fragment.
5. Locked final. Open an untouched segment once. If it is inspected and used to repair the model, it becomes development data for a new version.
6. Stress tests. Worsen execution costs, remove the best trade, and preserve clusters of wins and losses rather than shuffling them away.
7. Verdict. Assign an explicit status such as DIAGNOSTIC, REJECTED, DATA_NO_GO, or, only after sufficiently strong evidence, CONFIRMED.
These gates cannot guarantee profit. They perform a more modest and more useful job: they prevent weak evidence from receiving a strong name.
What a closed gate looks like
One carry experiment, V59, looked attractive during development. Its 126 development trades produced +0.4665% base expectancy per trade and +0.2665% under the frozen stress-cost model. Calibration remained positive across 35 trades: +0.3472% base and +0.1472% stressed.
Then the locked final was opened.
Only six trades occurred. Base expectancy remained positive at +0.1409%, but stressed expectancy became −0.0591%. Six observations cannot prove that the economic effect is impossible. They are enough to reject deployment under the precommitted rule because the required stressed final gate was not passed.
That distinction matters. “Not confirmed” is not the same claim as “disproved forever.” A negative gate closes one version; it does not license a universal conclusion.
Other branches stop even earlier. If the required historical source is missing, the point-in-time universe cannot be reconstructed without survivor bias, or the retrieved instrument violates the frozen source identity, then no profitability estimate is permitted. The honest result is DATA_NO_GO, not a backtest assembled from the nearest available substitute.
Why one p-value is not enough
A conventional p-value is easiest to interpret when the tested result was effectively alone. A selected trading result rarely is. Thresholds, windows, assets, prompts, filters, and repeated repairs create a family of alternatives.
We therefore record the search family and use robustness checks that answer different questions:
• cost stress: does the result survive less flattering execution economics?
• remove-best-trade: is the history carried by one fortunate event?
• block bootstrap: what happens when local clusters of wins and losses are preserved?
• locked evaluation: does the exact frozen rule survive data that did not help create it?
No single statistic replaces the others. A model can improve average return per retained trade while reducing total return or worsening drawdown. If the gate was defined as a joint economic condition, the researcher cannot replace it after the fact with whichever metric improved.
What we will publish
For each research branch, this series will report:
• the frozen hypothesis and failure rule;
• development, calibration, and locked-evaluation boundaries;
• the number of variants and trades when applicable;
• signal timing, execution timing, and cost assumptions;
• base and stressed outcomes;
• the evidence status and limits of the conclusion;
• negative results and data failures with the same precision as positive results.
We will not publish API keys, private endpoints, the complete production project, or parameter combinations that make a live edge mechanically reproducible. Internal reports are versioned and hashed; public links will appear only when a sanitized dossier can be opened without authorization and without exposing private research infrastructure.
The next QuantConnect note will examine a particularly tempting failure: a GPT veto filter improved some diagnostics, yet failed the full added-value gate, while the underlying mechanical family later produced negative base and stressed results on its locked final.
The laboratory rule is simple: the winner is not the strategy with the best chart, but the hypothesis that survives doors locked in advance.
────────────────────
Evidence-Based Trading Lab materials are for research and education. They are not investment advice, personalized financial advice, trading signals, or a guarantee of future returns. Historical results, simulations, and stress tests do not reproduce every condition of live trading.

Sergii
The material on this website is provided for informational purposes only and does not constitute an offer to sell, a solicitation to buy, or a recommendation or endorsement for any security or strategy, nor does it constitute an offer to provide investment advisory services by QuantConnect. In addition, the material offers no opinion with respect to the suitability of any security or specific investment. QuantConnect makes no guarantees as to the accuracy or completeness of the views expressed in the website. The views are subject to change, and may have become unreliable for various reasons, including changes in market conditions or economic circumstances. All investments involve risk, including loss of principal. You should consult with an investment professional before making any investment decisions.
To unlock posting to the community forums please complete at least 30% of Boot Camp.
You can continue your Boot Camp training progress from the terminal. We hope to see you in the community soon!