How every Backtestify report is produced

Engine backtestify-engine-4 · Parser v26

Signals and fills

A signal only counts once the candle that produced it has closed. The trade itself fills at the OPEN of the next candle, never on the signal candle's own close, which the strategy could not have traded on in real time.

Same-bar stop and target

When a stop and a target are both touched inside the same candle, the level nearer the candle's open is taken first, the same rule TradingView's tester applies, so the report and the exported script book the same exit. A candle that opens beyond a level fills at the open. Every such trade is counted and disclosed as ambiguous.

Gaps

When a candle opens beyond the stop, the fill is the OPEN price, not the stop price, the strategy could not have gotten the stop price because the market never traded there. The same rule applies in the other direction: a candle that gaps straight through the target fills at the open, not the target. Fills never happen at a price the candle did not actually trade.

Costs

Every trade pays commission on both sides (entry and exit) at 0.05% per side, plus 0.02% slippage per side. These are the same defaults every simulation uses, see the "Execution assumptions" panel on any report for the figures actually applied to it.

Position sizing and the notional cap

Each trade is sized off account risk to the stop distance. Where a strategy has no stop, or the implied position would be larger than the account can actually hold, size is capped at 100% of account equity notional, a single trade never simulates more capital than the account has.

Data

Candles are exchange spot data. When the primary data source cannot serve a symbol or window, Backtestify falls back to the next provider in a fixed order rather than inventing candles, a run that could not get real data reports a failure for that timeframe instead of a result.

Indicator warm-up

Every indicator needs bars before it produces its first value. Backtestify fetches 3× the longest moving-average period — and, for Wilder-smoothed indicators such as RSI, ATR and ADX, however many bars it takes for the seed to weigh under 0.1% — floored at 50 bars and capped at 1000 bars, purely to warm indicators up. Those bars are never counted as tested history and no trade is taken on them.

Train / test split

History is split chronologically: the first 70% trains or selects the strategy, the remaining 30% is the unseen window. In a backtest nothing is chosen on it. An improve run picks its shipped variant ON that window, among every candidate it measured there, and a create run compares its rules with yours there and can revert to yours: for both the window was read to choose the rules, so for them it is not out-of-sample evidence, whatever it shows. Robustness and cost-stress diagnostics re-run the same frozen spec on the same candles, parameter neighbourhood, execution delay, extra slippage, doubled costs, and those extra evaluations are reported as diagnostics, never as the verdict.

A "create" or "improve" run also reports how many candidate specs were evaluated on the train window; it is always shown on the report's integrity receipt.

The verdict

Two outcomes. RECOMMENDED means the run passes the owner-defined recommendation rule (verdict rules v2): every gate below, on the numbers the report shows. It does not prove the strategy is robust, and it is not advice.

Profit factor
≥ 1.1
Average trade
> 0
Net return
> 0
Max drawdown
≤ 40%
Trades, 1m – 2h
≥ 12
Trades, 4h and above
≥ 8
High win rate · ≥ 51%Large winners · PF ≥ 1.5Steady edge

Never recommended: a run that lost money, or one whose data or unseen window cannot carry the verdict (below). Rules we could not execute as written, no exit rule or a named strategy whose published rules we could not find are shown as warnings.

Since verdict rules v2 (26 September 2026) a new run is recommended only when every gate passes: it made money after costs, its data decided no trade through missing candles, its edge survives doubled costs, it has enough trades (below), its edge holds on the unseen window, and — when the result was the best of several timeframes or variants — that choice survives a significance test corrected for the number of tries. Reports made before then keep the verdict they were given and say “Verdict rules v1”.

The selection test, exactly (verdict rules v2). Null hypothesis: the mean per-trade result in R is ≤ 0 (without a stop, the per-trade net move). Statistic: one-sided t = mean / (sd / √n) over the trades the report shows, read against the normal distribution. α = 0.05 for the whole family. K = the timeframes a backtest evaluated; for an improve run that shipped a variant, the larger of that and the candidates measured on the window the variant was chosen on. Correction: Šidák, 1 − (1 − p)^K. A timeframe pays no penalty only when the request fixed it before any result was computed. Caveat: trades are treated as independent; serially correlated results make the test too lenient, and no block or Newey-West variance is applied.

Verdict rules v3 are defined and waiting for the owner to switch them on: Bonferroni (p × K), or Holm–Bonferroni when every candidate’s p is known; K = timeframes plus every candidate measured on the chosen window, marked “best of ≥ K” where earlier searches are not stored; and a run whose unseen window was read to choose its rules fails the unseen gate as “contaminated”.

A recommendation needs at least these trade counts (v1 reports only tagged them “limited sample”):

1m – 15m
100
30m – 1h
60
2h – 12h
40
1d and above
25

Unseen-window trades needed before its edge is scored:

1m – 15m
60
30m – 1h
40
2h – 12h
25
1d and above
15

What we do not claim

  • Every number is a hypothetical simulation of stated rules on past candles, it is not a live trading record and no capital was ever actually deployed.
  • No slippage or cost model survives all market conditions, especially thin liquidity or news events.
  • Results on one symbol or window do not transfer to another symbol, timeframe or period.

Public proof set

Real house-run reports, published with their full integrity receipt visible to signed-out visitors, including strategies that failed.