
Four levels of backtest honesty
An EMA test moved from +34% frictionless to -59% with its fullest cost and fill model. Review the assumptions, trade-count changes, and limits.
In this August 2026 experiment, a BTCUSDT EMA crossover on the 15-minute bull sample returned +34.24% with no trading costs and −59.21% under the study's fullest recorded cost and execution model. Its frictionless Sharpe ratio was 1.23, with 9.7% maximum drawdown. Under the full rung, Sharpe was −3.08 and maximum drawdown 60.5%.
The strategy rules, source candles, parameter values, and dates stayed fixed. The executions did not. Adding spread, slippage, finer intrabar resolution, and funding changed fill prices, exits, and even trade counts. This is a controlled comparison of four simulation configurations, not a claim that one fee line alone caused the entire gap.
The four rungs
We ran an EMA crossover (fast 12, slow 26) on BTCUSDT through four configurations. Each rung adds assumptions:
- Frictionless. No fees, no spread, no slippage. Orders use the rung's candle fill path.
- Fees only. A 5.5 basis-point taker fee per side, nothing else.
- Fees + spread + slippage. Add a 1 bp spread and 2 bps of market-order slippage, with 5 bps on stop fills.
- Full realism. Everything above, plus the intra-bar fill model (the engine replays 1-minute candles inside each bar to work out whether and when a stop or target was actually touched), slippage that scales with measured order-book conditions, and funding applied at recorded events.
Bybit's official execution documentation explains why market-order size can consume multiple book levels. That source describes exchange mechanics; the configuration below records how this study modelled them.
| Method field | Historical study record |
|---|---|
| Strategy | EMA(12, 26), BTCUSDT |
| Timeframes | 4-hour, 1-hour, and 15-minute |
| Bull sample | January 1, 2024 to June 30, 2025 |
| Bear sample | August 1, 2025 to August 1, 2026 |
| Fees-only rung | 5.5 bp taker fee per side |
| Spread/slippage rung | Fees plus 1 bp spread, 2 bp market slippage, and 5 bp stop slippage |
| Full rung | Above, plus 1-minute intrabar fills, dynamic slippage, L2 fills where covered, and funding |
| Exact venue and L2 coverage | Unknown in the result artifacts |
| Position-sizing rule | Unknown in the result artifacts |
| Exact engine commit | Unknown in the result artifacts |
The study recorded two contrasting historical windows: a bull window (January 2024 to June 2025, Bitcoin up 153%) and a bear window (August 2025 to July 2026, Bitcoin down 46%).
The bull window
| Timeframe | Frictionless | Fees only | + spread & slippage | Full realism |
|---|---|---|---|---|
| 4h | -4.46% | -8.99% | -10.33% | -19.05% |
| 1h | +28.57% | +10.27% | +0.88% | -13.61% |
| 15m | +34.24% | -27.06% | -50.70% | -59.21% |
Read the 1-hour row slowly, because it's the most instructive line here. The strategy starts at +28.57% with a Sharpe of 1.08, a result that would survive most people's sniff test. Charging fees alone takes it to +10.27%. Adding spread and slippage takes it to +0.88%, a rounding error away from breakeven. Only the last rung, the one that models when your stop actually got hit, drives it to -13.61% and a Sharpe of -0.48.
The strategy crosses from slightly positive to negative between rungs three and four in this sample. The result says the execution configuration matters; it does not guarantee the strategy will lose in every later sample.
The 15-minute row is the same direction with more trades. Sharpe goes 1.23, -1.17, -2.73, -3.08. Drawdown goes from 9.7% to 60.5%. The lowest-drawdown frictionless configuration became the highest-drawdown full configuration among these tested rows.
The bear window
| Timeframe | Frictionless | Fees only | + spread & slippage | Full realism |
|---|---|---|---|---|
| 4h | +1.72% | -2.15% | -4.91% | -5.43% |
| 1h | +3.34% | -14.71% | -15.24% | -22.30% |
| 15m | -13.48% | -35.79% | -40.76% | -56.62% |
Bitcoin fell 45.6% over this recorded window. Two of three timeframes are positive in the frictionless column, while all three are negative under the full model. This does not make the full rung “true” for every trader; its assumptions must be assessed against the intended venue and order size.
The recorded gap grew with trade count
Line up the gap between the frictionless and full-model results:
| Timeframe | Full-model trades | Frictionless → full-model gap |
|---|---|---|
| 4h | 97 | 14.6 points |
| 1h | 343 | 42.2 points |
| 15m | 1,479 | 93.5 points |
Within this experiment, the faster variants traded more and accumulated more cost. On the recorded $10,000 account, the 15-minute full-model row paid $3,632 in fees over the bull sample. The fees-only row recorded $4,549; its trade count and equity path differed, so this is not the cost of an identical trade list.
It also shows why “just trade a lower timeframe for more opportunities” is incomplete advice. More signals can mean more round trips, and every executed round trip has to overcome its own costs.
The part that surprised us
Look at the fees column across rungs two, three and four. They barely move. Fees land between 5.2 and 5.5 basis points of traded notional in every rung that charges them. Yet the results keep getting worse.
The additional difference in rungs three and four is not the explicit fee rate. Spread and slippage alter prices, while intrabar resolution can alter whether and when an exit occurs. The 15-minute bull rows recorded 1,551, 1,528, 1,471, and 1,479 trades across the four rungs. That count change shows that later rungs changed simulated executions, not merely the cash charge on a fixed trade list.
If a backtest models fees but fills at a signal price, its report should say so. For this strategy and sample, the execution-model changes were larger than the fee-only change; another market or strategy can have a different breakdown.
What to do with this
Record the assumptions before reading the result. If a report does not state fees, spread, slippage, funding, and fill timing, those fields are unknown. Do not silently substitute a “realistic” label.
Calculate the cost-adjusted margin. Ask how many basis points the strategy needs to earn per round trip before costs. At 11 bp of round-trip cost, a 20 bp average gross move retains 9 bp before other charges, while a 10 bp move is negative after that assumption. The inputs must match the intended execution.
Treat trade frequency as a cost decision, not a style preference. Before you drop to a faster timeframe, multiply the expected extra trades by your round-trip cost. That's the edge you now have to find just to stand still.
Resolve stops at a stated data frequency. A wide primary candle can contain both a stop and a target without revealing their order. Finer data can change the answer, but it cannot reconstruct events absent from the source feed.
Honest caveats
This is one strategy family on one asset and two historical windows. We chose an EMA crossover because its rules are simple and familiar, not because this study established a general verdict on crossover strategies. A strategy with a larger pre-cost margin may remain positive under the same assumptions; that has to be tested rather than inferred.
The exact venue, sizing rule, fee-schedule provenance, and L2 coverage were not preserved in the result artifacts, so “full realism” is the study's historical label rather than a universal benchmark. Identical EMA parameters ran on every timeframe and both windows; executions and trade counts changed as the model changed.
This experiment does not prove that no crossover can work or that costs always reverse a result. It shows one strategy whose sign reversed under its fullest recorded model, with a larger gap at the timeframes that traded more. Use the AlphaProve methodology and the individual spread, slippage, and funding definitions to specify the next test before comparing its numbers with these historical rows.