← Blog
Backtesting

What Should You Change First in a Crypto Backtest? 99.75 Million Tests

2026.07.22·12 min read·Rulyfi

Key takeaways

  • Neither entries nor exits controlled every result. Win rate was more sensitive to exits in all ten market-direction jobs, while maximum drawdown was more sensitive to entries in all ten.
  • Changing an entry indicator was a much larger move than nudging that indicator's period. Treating both as "entry tuning" hides the useful distinction.
  • For the study's held-out-data score, the specific entry-exit pair mattered most. Pairing-specific differences accounted for a median 90.8% of variation in the balanced comparison.
  • The practical order is outcome-specific: name the result you want to change, move the settings connected to it, then recheck the complete entry-exit cell.

When a backtest needs improvement, what should you change first: the entry signal or the take-profit and stop-loss rules?

It depends on the outcome. Win rate moved with exits. Maximum drawdown moved with entries. Total return was more entry-sensitive in seven of ten market-direction jobs. For the score measured on held-out periods, results depended mostly on the specific entry-exit pairing rather than either setting alone.

The useful question is not "entries or exits?" It is which result are you trying to change, and which setting actually moves it?

This study turns that question into a tuning map.

The experiment: 79,800 entries crossed with 125 exits

A contest between the single best entry and the single best exit would be biased. There are 79,800 entry rules here but only 125 exit rules. The larger entry pool gets far more chances to produce an extreme result.

So we did not compare champions. We measured how the full result population moved when each side of the strategy changed.

All numbers and charts in this article were calculated from a July 2026 Rulyfi Scanner research dataset.

ItemDesign
MarketsBTC, ETH, BNB, XRP, and SOL perpetual futures on Binance USDT-M
DirectionsLong and short evaluated separately, for 10 market-direction jobs
Window2021-05-28 08:00 UTC to 2026-07-21 13:00 UTC
Base timeframe1 hour
Entry indicators20 trend, mean-reversion, momentum, volatility, and volume indicators
Entry variations20 periods per indicator, 400 variations total
Entry rulesEvery distinct pair of the 400 variations, both conditions required: 79,800
Take profit1%, 2%, 4%, 8%, 16%
Stop loss1%, 2%, 4%, 8%, 16%
Time to live12, 48, 168, 504, 1,512 bars
Exit rules5 × 5 × 5 = 125
Results per job79,800 × 125 = 9,975,000
Total results9,975,000 × 10 = 99,750,000
Execution assumptionsMarket orders, 1x leverage, 0.04% taker fee and 0.01% slippage per fill, historical funding included
ValidationFirst half used in-sample; second half divided into four out-of-sample folds

The 20 indicators were SMA, HMA, SuperTrend, ADX, RSI, MFI, Williams %R, Z-score, CCI, ROC, Momentum, Fisher Transform, Bollinger Bands, Choppiness Index, R-squared, Percentile Rank, CMF, EFI, EMV, and Volume Ratio. They shared the period grid {5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 24, 28, 32, 40, 48, 56, 64, 80, 96}. We fixed the SuperTrend multiplier at 3 and the Bollinger Bands standard-deviation multiplier at 2.

TTL means maximum holding time, measured in base bars. On this one-hour scan, 12 bars is 12 hours and 1,512 bars is nine weeks. TP, SL, and TTL were always active together; the first condition reached closed the trade.

Return, maximum drawdown, win rate, and OOS analyses retained rows with at least 30 trades. That keeps a nearly inactive strategy from owning the tail of those distributions. Trade-count sensitivity used all rows because filtering on trade count would censor the outcome being measured.

The OOS score is a proxy, not an out-of-sample return

Our out-of-sample comparison uses walk-forward efficiency × ln(1 + total return). Walk-forward efficiency divides the average out-of-sample log-return rate by the in-sample log-return rate. Multiplying the two produces a score proportional to the out-of-sample log-return rate within each job.

We call it the OOS log-rate proxy score. It is not the percentage an account earned in the out-of-sample window. Among rows with at least 30 trades, 99.58% had a valid, unclamped proxy score.1

How we compared entry and exit sensitivity

Sensitivity here means how widely an outcome moves when a group of settings changes.

For the headline sensitivity analysis, we used the 76,000 entries that paired two different indicator families. We fixed one entry rule, took the median across its 125 exits, repeated that for all 76,000 cross-family entries, then measured the middle 50% range of those entry-level medians. The remaining 3,800 same-family pairs became a separate control.

For exit sensitivity, we reversed the calculation: fix one exit, take the median across the 76,000 cross-family entries, repeat across 125 exits, then measure the middle 50% range of the exit-level medians.

We divided both ranges by the middle 50% range of all results in that job. That normalization puts return, drawdown, win rate, and trade count on a comparable scale.

The effect ratio in this article is:

R = exit sensitivity ÷ entry sensitivity

  • R below 0.80 is entry-led
  • R above 1.25 is exit-led
  • Values in between are practical ties

With almost 100 million rows, tiny differences can easily look statistically significant. We did not use p-values to crown a winner. We used effect size and asked whether the classification recurred across the ten market-direction jobs. Those ten jobs, not 99,750,000 independent observations, are the recurrence unit.

One question produced five different answers

No side won every outcome.

OutcomeEntry sensitivityExit sensitivityEffect ratio REntry/tie/exit jobsReading
OOS log-rate proxy0.7090.3220.5326/3/1No universal winner
Total return0.7850.3440.3977/3/0Entry-led
Absolute max drawdown0.9390.3160.34110/0/0Entry-led
Win rate0.0491.00220.8140/0/10Exit-led
Trade count1.1230.4390.39410/0/0Entry-led

Entry versus exit sensitivity by outcomeRows show the OOS proxy score, total return, maximum drawdown, win rate, and trade count for the cross-family entry analysis. Columns show long and short jobs across five markets. Each cell is effect ratio R, exit sensitivity divided by entry sensitivity. Below 0.80 favors entry, above 1.25 favors exit, and the range between is a practical tie. Maximum drawdown and trade count favor entry in all ten jobs. Win rate favors exit in all ten.Entry versus exit sensitivity by outcomeR = exit sensitivity ÷ entry sensitivity, 76,000 cross-family entries per jobEntry-led, R < 0.80Practical tieExit-led, R > 1.25BTC LBTC SETH LETH SBNB LBNB SXRP LXRP SSOL LSOL SMedian Rentry/tie/exitOOS proxy0.830.370.350.471.620.410.400.860.590.880.536/3/1Total return0.620.300.310.350.940.380.360.940.410.960.407/3/0Max drawdown0.180.390.150.350.330.420.300.720.220.780.3410/0/0Win rate15.020.922.722.621.015.320.820.818.818.920.80/0/10Trade count0.430.410.390.400.360.390.420.400.390.380.3910/0/0L = long, S = short. Color shows the classification; every cell prints R.Each cell shows effect ratio R for the 76,000 cross-family entries, with exit sensitivity divided by entry sensitivity. We classified R below 0.80 as entry-led and R above 1.25 as exit-led. The median OOS proxy leans toward entry, but its 6/3/1 recurrence does not support a universal winner.

Win rate showed the cleanest division of labor. Exits led in all ten jobs, and their median effect was about 20.8 times the entry effect. If win rate is the outcome you want to move, checking TP and SL before hunting for another indicator is a defensible order.

Maximum drawdown ran the other way. Entries led in every job. Changing entry signals moved maximum drawdown more than the typical differences across exit combinations.

Total return also leaned toward entries, but not unanimously: seven entry-led jobs and three ties. The OOS proxy was weaker still. Its median ratio favored entries, but six entry-led jobs fell short of our prespecified seven-job recurrence threshold.

That might tempt us to summarize the study as "exits for win rate, entries for everything else." Even that is too coarse, because entry settings contain two very different kinds of change.

Changing the indicator is not the same as changing its period

Replacing RSI with SMA and moving RSI(14) to RSI(16) both get called entry tuning. They are not equivalent interventions.

We separated them. For an indicator change, the other entry leg and exit stayed fixed. For a period change, we moved one step along the period grid within the same indicator while holding everything else fixed.

The table reports the absolute matched difference divided by that job's overall middle 50% outcome range. Larger values mean one change moved the result more. They do not say whether the result improved.

These matched comparisons use the same 76,000-entry cross-family population as the headline analysis.

One setting changedOOS proxyTotal returnMax drawdownWin rateTrade count
Entry indicator0.5330.4270.3680.0370.280
Entry period, one step0.1630.1050.0800.0120.012
TP, one step0.2770.1850.1370.3640.096
SL, one step0.2930.1860.1610.3400.078
TTL, one step0.1300.0820.0570.0680.010

Normalized effect of changing one settingA heatmap of five changes and five outcomes. Values are not performance levels. Each is the absolute outcome difference from one change divided by that job’s middle fifty percent outcome range. Changing the entry indicator is largest for the OOS proxy, total return, maximum drawdown, and trade count. One-step TP and SL changes are largest for win rate.Normalized effect of changing one settingAbsolute matched change ÷ job outcome IQR, median across 10 jobsNot performance and not direction. Larger values mean the outcome moved more.OOS proxyTotal returnMax drawdownWin rateTrade countEntry indicator0.5330.4270.3680.0370.280Entry period +1 step0.1630.1050.0800.0120.012TP +1 step0.2770.1850.1370.3640.096SL +1 step0.2930.1860.1610.3400.078TTL +1 step0.1300.0820.0570.0680.010Smaller moveLarger moveNumbers in cells are exact.Each value is the absolute outcome difference between matched settings divided by that job’s outcome IQR. It measures magnitude, not performance level or improvement. Changing the entry indicator was the largest move for four outcomes; win rate reacted most to one-step TP and SL changes.

Changing the entry indicator was the largest move for the OOS proxy, total return, maximum drawdown, and trade count. Win rate reacted far more to TP and SL. One-step changes to entry period and TTL were small for most outcomes.

That changes the tuning order. Moving RSI from 14 to 16 before checking whether the broader indicator family fits the job uses a small knob to answer a structural question. Compare indicator families and exit shapes first. Refine periods inside a structure that has earned the extra attention.

A same-family control makes the distinction sharper. When we retained only the 3,800 entries that paired two periods from the same indicator, trade count changed from entry-led in all ten jobs to exit-led in nine and tied in one. Once indicator family was held constant, trade count became more exit-sensitive.

For the OOS proxy, the combination was the main event

Can you rank entries and exits separately, then join the best of each?

Not reliably in this result set.

We decomposed the complete factorial surface into the average entry effect, average exit effect, and their interaction. Interaction is the part that changes because a particular entry behaves differently under different exits. When interaction is large, the two main effects cannot describe the response on their own.2

Some OOS cells are unavailable because of the trade-count and validity filters. To keep the decomposition balanced, we rebuilt each cross-family matrix using only entries with a valid score under all 125 exits.

Across those complete matrices, interaction accounted for 79.5% to 95.7% of OOS proxy variation. The median was 90.8%. In all ten jobs, interaction was larger than either main effect.

OOS interaction share by market and directionTen horizontal bars start at zero. In the complete cross-family matrix restricted to entries valid under all 125 exits, the interaction share of OOS proxy variation ranges from 79.5 percent to 95.7 percent. The median is 90.8 percent.OOS interaction share by market and directionCross-family entries valid under all 125 exits only, n = 10 jobs0%25%50%75%100%BTC L89.6%BTC S94.4%ETH L89.9%ETH S90.7%BNB L90.9%BNB S89.1%XRP L79.5%XRP S95.7%SOL L95.1%SOL S95.1%Median R 90.8%Share left after entry-average and exit-average effects are removedBars start at zero. Combination-specific differences accounted for 79.5% to 95.7% of OOS proxy variation in the complete matrices, with a 90.8% median. Ranking entries and exits separately can therefore miss most of the response surface.

This is more specific than saying both entries and exits matter. Averaging one entry across 125 exits, or one exit across 76,000 cross-family entries, discarded most of the OOS response surface.

For OOS candidate selection, do not build separate entry and exit leaderboards and simply join their top rows. The final unit of judgment must be the combined entry-exit cell.

A tighter stop changed behavior, not the OOS direction

Stop width gives the most intuitive example.

Within the same cross-family population, comparing tight stops of 1% and 2% with wide stops of 8% and 16% produced a median job-level difference of 323.8 more trades and a 31.3 percentage-point lower win rate for the tight stops. That pattern is consistent with faster exits and more re-entry opportunities. We did not decompose trade paths by exit cause, so we do not claim that mechanism as proven.

The OOS proxy moved in different directions by market and by long or short. Trade count and win rate moved in consistent directions, but tighter stops did not produce a general OOS improvement.

That distinction matters. Raising win rate, reducing trade count, and improving an OOS result are separate jobs. If you try to make them all "better" without naming a primary outcome, the success criterion moves every time the settings do.

A practical tuning order from these results

The reusable result is not a favorite indicator or a magic TP. It is an order of operations.

1. Choose one target and its guardrails

If total return is the target, maximum drawdown and minimum trade count can serve as guardrails. If win rate is the target, keep total return and drawdown visible so a prettier hit rate cannot hide a damaged payoff structure.

2. Compare structural choices first

Cross distinct entry indicator families with broad TP and SL ranges. In this experiment, changing indicator family was a large axis for the OOS proxy, total return, drawdown, and trade count. TP and SL were the large axes for win rate.

At this stage, ask which structures move the target outcome at all. Do not spend the search on fine period choices yet.

3. Refine periods and TTL inside the chosen structure

One-step entry-period and TTL changes were generally smaller. Small does not mean useless. These are good axes for checking whether the neighborhood is smooth and avoiding a single brittle optimum after the larger structure is chosen.

4. Recheck entries and exits together

When OOS interaction is large, a "good entry" and a "good exit" may not exist independently. Check whether the selected entry survives the exit grid and whether the selected exit holds across neighboring entries.

Target outcomeCompare firstRefine laterFinal check
Win rateTP, SLTTL, entry periodTotal return and MDD together
Total returnEntry family, TP, SLEntry period, TTLEntry-exit cells by market and direction
Maximum drawdownEntry familySL, TP, periodReturn, win rate, and trade count together
Trade countEntry family, exit speedPeriod, TTLRepeat within a same-family control
OOS proxyEntry-exit combinationsPeriod neighborsAll 10 jobs and the complete exit grid

This is not a universal optimization formula. It is a starting order to recalculate on your own results. The important move is to map which settings move which outcomes before selecting the row that happens to rank first.

Four strong objections

Is the comparison fair when there are many more entries?

It would not be fair if we compared the best of 79,800 entries with the best of 125 exits. For the headline analysis, we instead took medians across the opposite block within the 76,000 cross-family entries and compared the middle 50% range of those medians. We also ran matched comparisons in which only one setting changed.

There is still a limit: indicator family is categorical, while the exit settings form numeric grids. They are not identical interventions. Our claims are therefore about observed sensitivity inside this fixed design, not causal superiority.

With 99.75 million results, doesn't everything become significant?

Yes, which is why statistical significance did not decide the headline. We used effect-size ratios and recurrence across ten market-direction jobs.

BTC, ETH, BNB, XRP, and SOL are not independent experimental subjects either. We describe the evidence as ten repeated cells within one crypto basket, not ten unrelated replications.

What if TP and SL are both touched in the same bar?

The main scan allowed an exit on the entry bar and resolved a same-bar TP-and-SL touch in favor of TP. That can be optimistic.

We replayed a representative sample with the conservative rule, choosing SL when both were touched. The dominant direction for maximum drawdown, win rate, and trade count remained stable in all ten jobs. For total return, XRP short moved from entry-led to a tie, and no job moved to exit-led.

The replay did not reconstruct the OOS folds, so it is not evidence for the OOS interaction result.

Does this transfer to other assets and timeframes?

Not directly. This experiment used five crypto perpetual futures, one-hour bars, automatic entry conditions, and a fixed TP, SL, and TTL grid. Daily equities or very short-term futures may produce a different sensitivity map.

What transfers is the measurement procedure: separate the outcomes, compare factor sensitivity in a common unit, and test whether interaction is larger than the main effects.

Build your own sensitivity map

The same method works on a smaller scan:

  1. Pick one outcome to improve and the guardrails it must respect.
  2. For entry-side sensitivity, take each complete entry rule's median across all exits, then measure the middle 50% range of those medians.
  3. For exit-side sensitivity, take each complete exit rule's median across all entries, then measure the middle 50% range of those medians.
  4. Normalize both ranges by the middle 50% range of all results for the same job and outcome.
  5. To measure an individual knob, compare matched rows where only that setting changes and everything else stays fixed.
  6. Confirm finalists as complete entry-exit cells across markets and directions.

Keeping the full result set makes this easier than working from a top-100 leaderboard. Rulyfi Scanner can expand entry combinations across a TP, SL, and TTL grid and save the complete results. The full result set lets you ask both which strategy ranked first and which setting moved the outcome you care about.

So what should you change first when tuning a backtest?

For win rate, start with TP and SL. For maximum drawdown and total return, start with the entry indicator family. Refine periods after choosing the larger structure. For the OOS proxy, judge entries and exits as combinations rather than separate winners.

A map connecting each knob to the result it can actually move turns tuning into a measurable order.

Frequently asked questions

Should I optimize entries or exits first?

There is no fixed order. In this experiment, win rate was more sensitive to TP and SL, while maximum drawdown and total return were more sensitive to entry indicator family. Choose the target metric first, compare its high-impact factors, then validate the final entry-exit combinations.

Should I stop optimizing entry periods?

No. Period changes were smaller than indicator-family and TP or SL changes, but they remain useful for checking stability around a chosen structure. The inefficient order is polishing one period before deciding whether the broader structure works.

What does 90.8% OOS interaction mean?

In the complete cross-family matrices containing only entries valid under all 125 exits, combination-specific differences accounted for a median 90.8% of OOS proxy variation after average entry and average exit effects were removed. Ranking entries and exits separately may therefore miss most of the surface.

Did this study find a profitable strategy?

No. It measured how settings moved total return, win rate, drawdown, trade count, and an OOS proxy. It does not certify any strategy's future profitability or live-trading suitability.


Trading, including automated trading, can result in loss of principal. This article is for educational purposes only, does not guarantee profit, and past backtest results do not predict future returns.

Footnotes

  1. Robert Pardo, The Evaluation and Optimization of Trading Strategies, Second Edition, Wiley, for practical background on walk-forward analysis and strategy optimization.

  2. NIST/SEMATECH e-Handbook of Statistical Methods, Estimate Main and Interaction Effects. We did not apply a classical two-level factorial design unchanged; we used robust summaries and a balanced-matrix decomposition suited to this large complete grid.

backtestingstrategy-optimizationentriesexitssensitivity-analysiswalk-forward