← Blog
Backtesting

We Ran 99.7 Million Crypto Backtests. One Market Produced All 103 Candidates.

2026.07.20·11 min read·Rulyfi

Key Takeaways

  • We ran 99,691,200 historical backtests across BTC, ETH, BNB, XRP, and SOL, with separate long and short searches. Of 69,140,072 strict eligible rows, 160 reached RMP 0.95. Every one was SOL long.
  • The blind primary DSR 0.95 screen passed zero rows. A post-analysis look through the pre-existing RMP gauge left 103 configurations after positive return and Sharpe, PSR of at least 0.95, positive walk-forward efficiency, and no more than one losing forward fold. They resolve to 33 exact indicator pairs and 40 distinct stored scorecards, so we do not treat them as 103 independent strategies.
  • The concentration had a fingerprint: 64 configurations paired two 4h Choppiness Index settings, and all 103 used 1.5% take profit with 1% stop loss. A representative 4h Choppiness(12) + Choppiness(96) row returned +1,732% historically over 1,022 trades, with a 55.68% win rate, -12.69% maximum drawdown, Sharpe 3.742, PSR 0.999, RMP 0.999, DSR 0 at N = 8,523,430 trials, walk-forward efficiency 0.129, and one losing forward fold out of four.
  • RMP is a search-sized random-maximum percentile, not a probability of profit. Rulyfi puts it beside PSR, DSR, and walk-forward evidence, then lets you save the entire result population and open a displayed candidate in Builder.

1. The entire map lit up in one corner

Five markets. Two directions. Three ways to pair the timeframes. The search had thirty market, direction, and timeframe cells where evidence could have concentrated.

Only one crossed the article's line.

The full population contained 99,691,200 backtest rows. A strict eligibility screen left 69,140,072 with at least 50 trades and finite, non-saturated statistics. Among those, 160 rows reached a Random-Max Percentile of 0.95. All 160 came from the SOL long job. Adding return, PSR, and walk-forward requirements left 103, still all SOL long.

The eligible populations were comparable rather than lopsided: the ten market-direction jobs retained between 6,777,270 and 7,006,322 strict rows each, and SOL long retained 6,899,691. The concentration therefore did not come from SOL receiving a materially larger eligible population.

The map below overlays two complete distributions on the same 0-to-1 vertical scale. Grey shows the PSR density and blue shows the RMP density for every strict row. The blue strip along the bottom includes 69,117,233 exact-zero RMP observations; the blue cells near the 0.95 line show where high-RMP evidence actually concentrated.

Overlaid trade-count density heatmaps for PSR and RMP by market and directionTen shared-scale panels arrange BTC, ETH, BNB, XRP, and SOL in columns and long and short directions in rows. Equal-size rectangular bins show trade count from zero to six thousand plus an overflow bin on the horizontal axis and metric values from zero to one on the vertical axis. Grey opacity is the global logarithmic density of all PSR observations. Blue opacity is the separately scaled global logarithmic density of all RMP observations, including exact zeros. The 160 RMP observations at or above 0.95 occur only in SOL long.Trade count × PSR and RMP density by market and direction69,140,072 strict rows · grey = PSR density · blue = RMP density · separate global log scalesBTCETHBNBXRPSOLLONG00.250.500.751.0160 RMP ≥ .95SHORT00.250.500.751.003k6k+03k6k+03k6k+03k6k+03k6k+RMP max 0.915.95.95PSR / RMP valuetrade count (0 to 6,000 linear · final bin ≥ 6,000)PSR densityRMP densityGrey cells count every PSR observation and blue cells count every RMP observation in matching bins of 150 trades by 0.025 metric value. Both layers contain all 69,140,072 strict rows and use separate global logarithmic density scales so their shapes remain comparable without treating PSR and RMP counts as one series. The blue bottom band includes 69,117,233 exact-zero RMP rows; the 160 RMP observations at or above 0.95 appear only in SOL long, and the article's forward screen leaves 103 configurations. The final trade-count column includes every row at 6,000 trades or more. Source: ten verified full exports, common window 2021-05-28 through 2026-07-18.

This is the first reason the result is more useful than a leaderboard champion. The maximum strict RMP outside SOL long was 0.915 on SOL short. Outside SOL entirely, the highest market and direction was XRP short at 0.196. BTC long peaked at 0.0014; the remaining BTC, ETH, and BNB cells were effectively zero at the displayed precision.

The finding is a location in a search, not a forecast. If we had opened only the top-return row from each market, the concentration would have been invisible. It appeared because the scanner kept a search-sized statistic on every row and the full export kept every row available for comparison.

2. Five markets, two directions, one fixed search

The campaign began with a different question. We wanted to know whether a higher-timeframe context paired with a lower-timeframe trigger would produce strategy neighborhoods that repeated across major crypto markets. Same-timeframe pairs were included as controls, and long and short were split before the results were read.

The data rejected the strong version of that idea. Only 9 of the 103 final configurations were cross-timeframe. The largest group, 64 rows, paired two indicators on 4h. That is a useful property of a scan: the original hypothesis sets the grid, but it does not get to choose the answer.

Here is the campaign receipt.

MarketsBTC, ETH, BNB, XRP, SOL perpetual futures on Binance USDT-M
Window2021-05-28 08:00 UTC to 2026-07-18 15:00 UTC, 45,055 hourly base bars
DirectionsSeparate long and short jobs for every market
Indicator grid322 variants on 4h and 323 variants on 1h
PairingEvery two-variant combination, split into same-4h, same-1h, and 4h + 1h
ExitsTP {1.5, 3, 6, 12}% by SL {1, 2, 4}% by TTL {12, 36, 84, 168} hours, 48 exits
Rows9,969,120 per job, 10 jobs, 99,691,200 total
CostsMarket orders, 0.04% taker fee and 0.01% slippage per fill, 1x leverage
FuturesHistorical funding included; positive funding costs longs and credits shorts
ValidationFirst half of the window as in-sample; second half split into four equal-time forward folds; minimum 50 trades; full result export enabled

"645 indicator variants" deserves the same specificity as the result count. In the table below, P4 is {8, 10, 12, 14, 16, 20, 24, 28, 32, 36, 40, 48, 56, 64, 72, 80, 96, 112, 144, 192} and P1 is {3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 28, 32, 40, 48, 56, 64, 72, 88, 112}. Here is the exact indicator-by-indicator sweep.

TimeframeIndicator groupExact sweepVariants
4hSMA, EMA, HMA, VWMA, ADX, Aroon, Vortex, Donchian, Choppiness, R-squaredperiod P4200
4hHurst Exponentperiod {32, 36, 40, 44, 48, 52, 56, 60, 64, 72, 80, 96, 112, 132, 160, 192, 224, 256, 288, 320, 352, 384, 448, 512}24
4hKAMAperiod {8, 10, 14, 20, 28, 40, 56, 80, 112, 160}; fast 2; slow 3010
4hSuperTrendperiod {7, 10, 14, 20, 28, 40, 56, 80}; multiplier {1.5, 2, 2.5, 3, 4}40
4hMACDfast {5, 8, 12, 16}; slow {24, 32, 48, 64}; signal {5, 9, 13}48
1hRSI, Stochastic, MFI, Williams %R, CMO, Percentile Rank, Z-scoreperiod P1; Stochastic D period 3168
1hStochastic RSI, Connors RSI, Ultimate Oscillator, MACD, TSI, TRIX, Bollinger BandsStochastic RSI {7,14,21} × {7,14,21} × K {2,3}, D 3; Connors {2,3,5} × {2,3} × {20,50,100}; Ultimate {3,5,7} × {14,21,28} × {56,84}; MACD {4,8,12} × {24,36,48,72} × {5,9}; TSI {20,30,45} × {5,8,12,16} × {5,9}; TRIX {6,9,12,18,24,48} × signal {3,5,7,9}; Bollinger period {10,14,20,28,40,56,112} × deviation {1.5,2,3}147
1hHammer, Engulfing, Morning Star, Three White Soldiers, Harami, Piercing Line, Marubozu, Tweezer Top/Bottomno parameters8

That is 322 variants on 4h and 323 on 1h, 645 in total. Entry conditions were the scanner's automatic defaults for each indicator. MACD, TSI, and TRIX used their crossover or crossunder presets, and both conditions in every pair were required.

The results assume the stated flat fee and slippage on every fill. They do not simulate order-book depth, partial fills, or market impact.

The representative Choppiness pair is also mechanically simple. Each Choppiness Index is a direction-neutral trend filter that must read below 38.2, and both the 12-period and 96-period filters must be true on the 4h bar. The separate long job supplies the direction; the two indicators identify a trending regime rather than choosing long on their own.

Each job expanded 645 indicator variants into 207,690 distinct pairs, then tested all 48 exits. Display labels were not used to classify the rows because they shorten strategy definitions. The analysis decoded each pair from its underlying variant indices, which preserves its timeframe, indicator, and exact parameters.

That fixed receipt is what turns the headline from a sorted anecdote into a population result. Every market received the same grid, the same common window, the same exit choices, and the same evidence rules.

3. How 160 became 103

RMP asks one question: under its model, what is the probability that the best result produced by pure chance across a search of this size would fall below this row? It uses the row's trade evidence and the scan's recorded trial count. An RMP of 0.95 is a search-aware hurdle. It is not a 95% probability that the strategy will make money. The full definition, calibration, and assumptions are in our RMP study.1

Three population counts appear in this article because they answer different questions. Each SOL long job generated 9,969,120 configurations. Of those, 8,523,430 had a finite per-trade Sharpe and entered the search-wide DSR and RMP trial population. The stricter article eligibility screen later left 6,899,691 SOL long rows with at least 50 trades and non-saturated statistics. N therefore means statistical trials, not generated rows or final eligible rows.

RMP and its 0.95 interpretation were published and shipped before this campaign, but RMP was not the campaign's original headline test. A blind post-export DSR 0.95 screen passed zero rows. This article is a post-analysis exploration of the pre-existing supplemental RMP gauge, which is why it calls the 103 rows candidates and reserves confirmation for a frozen future test.

Before RMP entered the screen, a row had to pass the mechanical eligibility rules:

  • at least 50 realized trades;
  • finite Sharpe, t-statistic, per-trade Sharpe, payoff, and walk-forward efficiency;
  • no engine saturation or sentinel value in those fields.

That left 69,140,072 of the 99,691,200 rows. RMP of at least 0.95 then left 160. Their timeframe split was already uneven: 121 same-4h, 30 same-1h, and 9 cross-timeframe. Every row was SOL long.

The article screen then required:

  • total return above zero;
  • Sharpe above zero;
  • PSR of at least 0.95;
  • walk-forward efficiency above zero;
  • no more than one negative fold among the four forward folds.

The first four conditions removed none of the 160. The negative-fold requirement removed 57 same-4h rows, leaving 103. That detail is important. The final funnel did not manufacture the SOL concentration by cutting away other markets. The concentration was complete at the RMP stage, before the forward-fold condition removed anything.

It also tells us what the 103 means. These configurations had positive historical returns, strong single-row significance, adjusted evidence at or above the 95th percentile of the modeled search-sized random maximum, and positive forward efficiency with at least three of four forward folds non-negative. They are candidates rather than certified future edges. The useful output is a compact set worth comparing and stress-testing.

4. The 103 formed neighborhoods

One exceptional row can be a narrow accident. A useful scanner result should make it possible to look around the row: nearby periods, related indicator families, alternative exits, and repeated scorecards.

The 103 configurations were not evenly scattered.

Composition of the 103 SOL-long candidate configurationsA stacked bar divides 103 configurations into 64 same-four-hour, 30 same-one-hour, and 9 cross-timeframe rows. A horizontal bar chart divides the same 103 rows into 64 Choppiness plus Choppiness, 14 RSI plus Stochastic, 14 RSI plus Williams R, 9 Hurst plus RSI, and 2 RSI plus Percentile Rank configurations. The rows collapse to 33 exact indicator pairs and 40 distinct stored scorecards. Every row uses 1.5 percent take profit and 1 percent stop loss.Composition of the 103 candidate configurationsCounts from the SOL-long article screen · one row can repeat an equivalent stored scorecardTIMEFRAME PAIR64same 4h30same 1h94h + 1hINDICATOR FAMILY016324864Choppiness + Choppiness64RSI + Stochastic14RSI + Williams R14Hurst + RSI9RSI + Percentile Rank2103 configurations · 33 exact indicator pairs · 40 stored scorecards · TP 1.5% / SL 1.0% in all 103 rowsThe two views use the same denominator of 103 configurations. The top bar shows the timeframe composition; the lower bars show the indicator-family counts. The candidate population is concentrated rather than uniform: 64 rows pair two 4h Choppiness settings. Duplicate checks reduce the population to 33 exact pairs and 40 stored scorecards. Source: SOL long full export and deterministic variant decoding.

The biggest block is exact. All 64 same-4h rows pair two Choppiness Index variants. They cover 17 exact period pairs. Most combine a short period around 10 to 14 with a longer period from 72 to 112. The same-1h group contains 30 rows, led by RSI(24) paired with Stochastic or Williams %R periods around 56 to 72. The nine cross-timeframe rows pair a long-window 4h Hurst estimate with a 1h RSI, mostly RSI(24).

Every one of the 103 rows uses a 1.5% take profit and a 1% stop loss. The TTL varies, but not every TTL produces a new economic result. Collapsing rows by exact indicator labels leaves 33 pairs. Collapsing rows with identical stored return, win rate, drawdown, trade count, Sharpe, PSR, DSR, RMP, walk-forward efficiency, and negative-fold count leaves 40 distinct scorecards.

That distinction prevents a common research mistake. The row count is 103, but it is not evidence from 103 independent strategies. Parameter neighbors share signals and trades. TTL settings can also become irrelevant when another exit closes every trade first. The useful object is the shape of the cluster and the scorecards inside it.

The Choppiness block makes that shape visible.

The 17 candidate Choppiness period pairs on SOL longOccupancy matrix with the shorter four-hour Choppiness period on rows and the longer period on columns. Seventeen cells contain candidates and each occupied cell shows its TTL row count. The period 12 with period 96 cell is outlined and its maximum RMP, display-capped at 0.999, appears once in the representative evidence rail.17EXACT PERIOD PAIRSinside 64 Choppiness rowsLONGER 4H CHOPPINESS PERIOD40485664728096112shorter 4h period84 rows101 row3 rows4 rows4 rows4 rows4 rows124 rows4 rows4 rows4 rows4 rows4 rows4 rows144 rows4 rows204 rowsREPRESENTATIVE PAIR12 + 96maximum RMP0.9994 TTL rowsone stored scorecardoutlined in the matrixEach filled cell is an exact pair of two 4h Choppiness Index periods among the article candidates; its label gives only the number of TTL rows. All 17 cells already cleared the article's RMP gate, so the matrix uses one occupancy color instead of exaggerating tiny differences near 1. The outlined 12 and 96 pair has four TTL rows that collapse to one stored scorecard; its maximum RMP appears once in the evidence rail, capped at 0.999 for display. Source: SOL long full export, 17 exact pairs and 64 rows.

The outlined 12 and 96 pair owns the highest RMP in the search, but it is not alone. Period 12 also appears with 40, 56, 64, 72, 80, and 112. Period 10 appears with 40, 48, 72, 80, 96, and 112. Periods 8, 14, and 20 contribute additional cells around the same longer-period region.

This is the product moment hidden by a normal top-100 table. The interesting result is not simply a rank-one configuration. It is a rank-one configuration with company, visible only when the population is saved and its neighborhood is inspected.

5. One row, every metric

The 4h Choppiness(12) + Choppiness(96) pair is a useful representative because it has the search's maximum RMP and four TTL rows with identical stored scorecards. The complete row reads:

FieldResult
DirectionSOL long
Take profit / stop loss1.5% / 1.0%
TTLs with identical scorecard12, 36, 84, 168 hours
Trades1,022
Total return+1,732.01%
Win rate55.68%
Maximum drawdown-12.69%
Sharpe3.742
PSR0.999
DSR (N = 8,523,430 trials)0.000
RMP0.999
Walk-forward efficiency0.129
Negative forward folds1 of 4

Total return is compounded trade capital over the full five-year window under the campaign's fill and cost assumptions. The number is large, but it is not an isolated outlier inside the candidate set. Across all 103 configurations, the median return was +1,115.94%, median maximum drawdown was -12.60%, median win rate was 54.90%, and median trade count was 980. Returns ranged from +229.69% to +1,760.06%; drawdowns ranged from -26.22% to -5.89%; trade counts ranged from 251 to 1,553.

The validation fields keep the attractive numbers in context. The display caps both near-one values at 0.999, but PSR and RMP answer different questions and neither means certainty. PSR says the row's per-trade record is extremely unlikely to have a Sharpe at or below zero under the model. RMP places the row essentially at the top of the modeled search-sized random-maximum scale. DSR is exactly zero. Walk-forward efficiency of 0.129 says the mean out-of-sample log-return rate was about 12.9% of the in-sample rate, positive but far below full carry-over; one of the four forward folds lost money.

A return-only leaderboard would display the 1,732% and stop. The Scanner result detail keeps the disagreement beside it, which is exactly where a research decision begins.

6. Why RMP and DSR disagree

PSR, DSR, and RMP do not vote on the same proposition.

MetricQuestion
PSRIs this row's Sharpe above zero, viewed on its own?
DSRIs its per-trade Sharpe above the expected best Sharpe in this search?
RMPWhat percentile does its adjusted trade evidence reach against the modeled search-sized random maximum?

DSR comes from the Deflated Sharpe Ratio and corrects for selection across many trials.2 It expresses its hurdle in per-trade Sharpe units. In mixed searches, low-trade rows can make that Sharpe distribution wide and set a high expected-maximum bar. A busy strategy may accumulate strong whole-period evidence while its per-trade Sharpe stays below that bar.

RMP was added for that high-trade-count region. It calculates the probability that the modeled maximum across the recorded trials would fall below the row's adjusted t-statistic, which includes the square root of trade count. That is how a 1,022-trade row can read DSR 0 and a display-capped RMP of 0.999 without either calculation being missing or broken.

The whole candidate population shows the same geometry. All 103 rows come from the same N = 8,523,430-trial SOL long search and have DSR 0. Their trade counts range from 251 to 1,553, with a median of 980. Displayed RMP ranges from 0.951 to 0.999, with near-one values capped at three decimals. The two columns preserve different questions: DSR rejects the per-trade Sharpe against its search-wide hurdle, while RMP places the accumulated adjusted evidence near the top of its modeled scale. Because RMP was examined after the primary DSR screen returned zero rows, these remain discovery candidates for a frozen follow-up test.

That is why Rulyfi ships the three gauges together. A single score would force one geometry onto every strategy frequency. The useful result screen lets the row remain attractive and questionable at the same time.

7. What the concentration establishes

The result establishes a narrow historical fact: within this fixed grid, common window, cost model, and five-market basket, every strict RMP 0.95 row came from SOL long. After the article's forward screen, the remaining configurations concentrated in a small set of indicator families and one shared TP/SL combination, while TTL still varied.

It does not establish that Solana will rise, that a long position should be opened, or that the 103 rows are independent confirmations. All five crypto markets share regimes and liquidity shocks. The candidates were selected and evaluated on the same historical campaign, so this is a discovery set rather than a fresh test of a frozen hypothesis.

Execution assumptions also matter. Funding is included, but order-book depth and partial fills are not. The flat 0.01% slippage may be too low for some real order sizes.

The result is still actionable as research. It tells us exactly what to freeze next: SOL long, the visible Choppiness and RSI neighborhoods, the 1.5%/1% exit shape, and the untouched future period that arrives after this search. The full population narrowed a huge design space to a compact set that can be challenged without pretending the challenge is already won.

8. Run the neighborhood test on your own scan

You do not need a 99.7 million-row campaign to use the same workflow.

  1. Start with a market and hypothesis you actually care about. Give each indicator a range broad enough to expose neighbors rather than one hand-picked period.
  2. Run the scan with a fixed exit grid and switch on Save all results before launch. Download the file when the scan finishes because generated exports remain available for seven hours. If every result came from cache, choose Re-run without cache (uses credits) in the download panel to create a fresh export.
  3. When validation is enabled, read PSR, DSR, RMP, efficiency, and negative folds beside return, win rate, drawdown, and trade count.
  4. Save all results. Check whether the top row is isolated, repeated only by irrelevant TTL changes, or supported by nearby indicator parameters and exits.
  5. Open a displayed candidate in Builder. Change one assumption at a time: fill convention, costs, period boundary, or the next untouched window. Candidates found only in the downloaded file must currently be recreated in Builder rather than opened directly.

The campaign in this article began with a cross-timeframe idea and ended with a same-4h Choppiness concentration. That reversal is the reason to scan a population instead of tuning a favorite row. The question worth carrying into your next search is simple: does your best result have company?

Frequently asked questions

Is an RMP of 0.95 a 95% probability of making money?

No. It is the probability, under RMP's model, that the random maximum from a search of this size would land below the row's adjusted evidence score. Here it marks a high modeled percentile on one historical window. Because this article examined RMP after the primary DSR screen returned zero rows, it identifies evidence for a frozen follow-up test, not proof that search luck has been eliminated or that the strategy will profit.

Why is DSR zero at N = 8,523,430 when RMP is nearly one?

DSR grades per-trade Sharpe against the expected best per-trade Sharpe in the search. RMP reports where adjusted trade evidence, including trade count, sits against the modeled distribution of a search-sized random maximum. High-trade-count rows can accumulate strong evidence while their per-trade Sharpe remains below DSR's hurdle. The RMP article shows the geometry and calibration in detail.

Does this mean I should go long Solana?

No. This is a historical concentration inside one fixed search, not a live signal or forecast. The useful output is a compact research neighborhood whose costs, fills, windows, and future performance can now be tested deliberately.

How do I check whether my own best result has neighbors?

Run a parameter range rather than one setting, switch on Save all results, and compare nearby periods and exits in the export. Keep return, win rate, maximum drawdown, and trade count beside PSR, DSR, RMP, and walk-forward evidence. Open a displayed candidate in Builder; recreate an export-only configuration there if you want to change its assumptions.


Auto-trading and trading carry a risk of losing your principal. This article is educational, does not guarantee profit, and past backtest results do not predict future returns.

Footnotes

  1. Šidák, Z. (1967). "Rectangular Confidence Regions for the Means of Multivariate Normal Distributions." Journal of the American Statistical Association, 62(318). RMP uses the corresponding one-sided multiple-testing correction in percentile form and remains an exploratory approximation with assumptions described in the linked RMP study.

  2. Bailey, D. H., & López de Prado, M. (2014). "The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality". The Journal of Portfolio Management, 40(5).

solanacrypto-backtestingrandom-max-percentilewalk-forwardstrategy-scanner