TFF · Observed behavior
Examine the actual trading sequence: sizes, follow-up purchases, costs, holding period and open risks. Data authenticity and completeness remain upstream questions.
Tradeflow ForensicsStrategies · methodology · evidence review
Understanding simulation, recognizing misconceptions and retracing the path from historical explanation to prospective testing.
A backtest applies a defined set of trading rules to historical data under specified assumptions. It asks: How would this version of the strategy have performed in this dataset and execution model? It shows no actual account performance achieved and alone does not provide a reliable promise for future profits.
This makes it valuable: rules can be controlled, mistakes discovered, costs examined and unsuitable ideas discarded before real capital is used. The transition from a retrospective calculation to a forward-looking assertion becomes problematic if its presuppositions are concealed.
“Data-based” initially only means that data has been used. Whether the data were available in a timely manner, the rule was only created after consideration of the result and the execution was realistically modeled, remains open.
| Retrospective investigation | Decision in the current market |
|---|---|
| The later course is known to the researcher. | The further course is unknown. |
| Rules, time period and parameters can be changed repeatedly. | The decision must be made with the available knowledge. |
| The best variant can be selected later. | Which variant fits in the future is still open. |
| Execution and costs are modelled. | Requests meet real quotes, liquidity and operating conditions. |
| Failed attempts can disappear from the presentation. | Wrong decisions cause real costs and capital losses. |
A retrospective analysis does not necessarily use future information. A sound test reconstructs what was known at each point and generates decisions in chronological order. For example, a signal based on the daily closing price cannot also be filled at an earlier intraday price if the information required to generate it was only available later.
The difference still remains: the researcher can know the course, even if the program code only reads past prices. Knowledge of a later crisis can already influence the choice of instrument, filter or test period. Therefore, “no future access in the code” is not sufficient as proof against all retrospective distortions.
Predictive validity requires additional assumptions: The context under investigation must continue, the available data must be comparable and the strategy must remain actually executable. An interest rate regime, market structure, competitor or broker conditions may change. The future is not another line of the same immutable experimental setup.
A screenshot does not reconstruct data quality or selection process. A second evaluation of the same file checks calculations, but does not automatically confirm the authenticity or completeness of the initial data. A file hash protects the assignment of a version; it does not prove that the trading data is true.
Among other things, MT5 distinguishes real ticks, generated ticks, 1-minute OHLC and open-prices tests. Suitability depends on whether a strategy responds to movements within a candle. A test only on opening courses cannot test an intrabar-dependent logic equally. The tester offers deceleration models; the mode without deceleration expressly maps idealized conditions. In the case of profit calculation in pips, swap, commission and margin audit are eliminated according to the documentation. [1]
Real ticks are closer to the observed course, not a guarantee of completeness. The MT5 documentation also describes generated ticks as a replacement when tick data is missing from an existing minute candle. For real ticks, the spread can change within the minute; for generated ticks, the corresponding minute spread is used. In addition, a last chart does not replace the bid/ask prices that are relevant for trading. [2]
High modeling quality in itself says nothing about economic viability, independent choice or future profitability. Even good price data does not automatically know your real position in an order queue, every rejected order ticket or the market impact of a large position.
MT4 and MT5 should not be equated via seemingly identical report labels. Test data, models and trade mechanics. The MT4-/MT5-Strategy Tester GuideBacktests and Track Records Describes the operation; this page assesses the significance. Additional: Execution quality, Contract data and Cost models.
| Vulnerability | Mechanism | Control question |
|---|---|---|
| Look-ahead / Data leak | Later known prices, revised data or future characteristics flow into earlier decisions. | When was all the information actually available? |
| Overfitting | The rule adapts to particularities of the known data set. | Does performance persist with neighboring parameters and new data? |
| Data snooping / multiple selection | Many variants are tested; the best hit is shown alone. | How many attempts preceded the version shown? |
| Survivorship bias | Failed instruments, accounts or strategies are missing. | What elements were available at the time and were later removed? |
| Cherry picking | Favorable periods, symbols or key figures are highlighted. | What does the pre-determined full evaluation look like? |
| Unrealistic fills | Any desired price is considered available, regardless of order type and market conditions. | What assumption creates the fill and when does it fail? |
| Hidden open loss | Balance contains closed gains while losing positions continue to run. | What is the worst equity value? |
| Incomplete costs | Spread, fees, financing or slippage are missing or too cheap. | Does a result remain after realistic costs? |
Bailey, Borwein, López de Prado and Zhu are investigating selection bias through backtest optimisation. Their approach considers, among other things, the relationship between in-sample selection and out-of-sample performance. A simple withheld dataset does not automatically answer how much previous multiple attempts affected the selection. Their research is not a blanket diagnosis of each EA; it justifies why the entire selection process must be part of the proof. [3]
Robust results don’t have to look spectacular. A wide range of similar results for neighbouring parameters should be assessed differently than a single narrow peak value. Even such an area is not proof of the future: the entire parameter family can depend on the same historical market phase.
All candidates in this teaching model have the same expected gross value of zero. Each has 60 randomly produced historical results and 60 other independent results. The selection is based solely on the highest historical total profit. The second half does not affect the selection.
Own synthetic simulation: per step independently +10 or −10 money with equal probability; constant size, no costs, no real rates. The starting value makes the experiment reproducible. For the same starting value, smaller candidate sets are subsets of larger sets. No testing of real strategies, no PBO calculation, and no probability of future success.
The selected candidate can win or lose in the second section. Both are possible. A single further positive result therefore proves no advantage. Consciously change the starting value and compare several complete experiments. Keeping only pleasant experiments would re-create the same selection error.
An invented account starts with 10,000 money. It closes a gross profit of 100 in twelve steps each, while another position temporarily carries an open loss of 2,000. The lines show the complete artificial course, including the starting value.
No deposits/payments, no interest and no margin calculation. The costs are a flat deduction per closed trade and are not intended to replicate a brokerage fee. Open P/L is additionally taken into account without deducting the costs again. Drawdown is calculated from the displayed model times including starting value; unobserved intermediate movements are missing.
That the same strategy could produce other curves is not shown by this fixed example. This would require further paths and verifiable assumptions. The laboratory demonstrates the effect of lack of information on the observation of a given course.
| Characteristic | Calculation / Significance | Typical fallback |
|---|---|---|
| Hit rate | Winning closed trades / all considered closed trades. | Many small gains do not make rare large losses unimportant. |
| Profit factor | Total positive / amount of the sum of negative trade results. | Without losing trade, the quotient is not finally determinable; no freedom from risk. |
| Historical expected value | Sum of trade results/number considered. | A sample average is not automatically the expected value of the next trade. |
| Equity drawdown | Decline from a previous equity high to a later value. | The historical maximum is not an upper limit for future losses. |
| Sharpe ratio | Excess yield in relation to yield fluctuation; indicate frequency and method of calculation. | A high value does not guarantee freedom from loss or low extreme risks. |
| Exposure/holding period | Tied positions and duration of risk exposure. | A return figure alone does not show what risk was taken for it. |
The MT5 test report defines profit factor, various balance/equity drawdowns and Sharpe ratio. The concrete implementations and reference values must be considered in the comparison. [4] Measurements should be considered net, over comparable periods and with reported open positions.
Deposits are not a trading profit. With changing capital inflows, simple change in account value, time-weighted return and money-weighted return differ. Even the annualization of short test periods can create a false impression of stability. An accurately calculated number cannot improve a weak data base.
Definition limit in MT5: The “forward” section of the strategy tester is a later part of the selected historical period. It is not automatically a future live or demo period. Its independence depends on whether it has already been incorporated into the development. [1]
None of these methods alone provide a seal of quality. For rare extreme events, even a long profitable phase may contain few observations. Many closely related trades are not as much independent evidence. A rigid minimum number or universal drawdown boundary would be misleading without a strategy and data context.
| Claims | What is actually proven | Missing test |
|---|---|---|
| “Ten years profitable.” | A certain historical test can be positive. | Rules selected in advance or afterwards? Total cost? |
| “99% data quality.” | A technical quality indicator can describe a data aspect. | No proof of a 99 percent chance of winning. |
| “AI recognizes the next market regime.” | One model was studied under certain data conditions. | Data leak, selection process, independent forecasting performance. |
| “Maximum 8% drawdown.” | A measured historical decline in a particular series. | No fixed loss cap for later paths. |
| “Independently verified.” | A third party has checked an explicitly designated test area. | What exactly: origin, calculation, completeness or account access? |
| “Forward Confirmed.” | A further test series was evaluated. | Was it unknown and unchanged before the beginning? |
A methodological error can arise from ignorance. A shortened presentation may be negligent. If unsuccessful variants, open losses or unrealistic assumptions are deliberately concealed, the presentation can create a wrong impression. The intention cannot be determined from the curve alone. Verifiable data and omissions are therefore critically examined.
As an example of this distinction, the US-SEC treats backtest results in the context of its rules for investment advisor marketing as hypothetical performance. Their materials emphasize the importance of assumptions, risks and limitations. This is a US-specific regulatory context and not a blanket legal statement for every EA seller in Europe. [5]
Download the checklist (markdown)
Ask providers for the complete test protocol, strategy version and parameters, data source, all cost assumptions, the selection process, all open positions and a traceable independent test phase. Ask for changes between backtest and live account, and full transaction data instead of screenshots only.
Decision rule: A lack of proof is not evidence of fraud. However, it limits what conclusions you can responsibly draw. A spectacular result without a reproducible basis deserves less trust than a more modest, cleanly documented investigation.
Examine the actual trading sequence: sizes, follow-up purchases, costs, holding period and open risks. Data authenticity and completeness remain upstream questions.
Tradeflow ForensicsView altered orders and conditions. The model under study determines the validity; a simulation is not a validation of its own assumptions.
Grid and margin laboratoryFrom proof of strategy to user account: size, timing, cost and execution can change the result.
Social and Copy Trading Knowledge · copytrader.io ↗The references combine technical verification tasks; no implemented data integration is claimed. Additional: Exporting reports, Understanding Expert Advisors and Secure evidence and protect consumers.
Research status 01.10.2026. Manufacturer sources explain tester functions and report definitions; the scientific work examines selection bias. The critical classification and calculation examples are our analytical workup. The teaching models do not use real market data and do not evaluate a specific product.