Funding Barely Changed Average Returns. It Changed 10.8% of Backtest Candidates
Key takeaways
- Turning funding on changed overall average total return by only -0.0083 percentage points. Positive and negative effects across ten market-direction groups largely canceled one another out.
- Under a candidate rule of at least 30 trades and positive total return, 162,165 decisions changed. That was 10.8% of the 1,501,027-setting candidate union, or about one in nine.
- Across ten Top 100 lists, 134 of 1,000 slots changed. Six of the ten lists still retained at least 90% of their members, and most rank movement across the full populations was small.
- The practical answer is not to subtract a fixed funding cost. Run the same scan with funding OFF and ON, then compare whether the candidate list, ranking, and validation evidence hold up together.
Does adding funding to a backtest merely nudge returns, or can it change which candidates advance to the next round of validation?
In a Rulyfi Scanner study, we matched 48,825,000 pairs of results with historical funding turned off and on while keeping each setting identical. Overall average total return changed by -0.0083 percentage points. That number alone might suggest funding was safe to ignore. Yet 162,165 candidate decisions changed, and 134 of 1,000 slots were replaced across ten separate market-direction Top 100 lists.
The average did not show that funding had no effect. The sign varied by market and direction, and those differences canceled out inside a single number.
Method: change only Funding fees for the same setting
The comparison did not match attractive results after the fact. We held the entry signals, exit rules, candles, cost assumptions, calculator version, and the fill rule that gives SL priority when TP and SL are both touched within the same candle constant. The only change was setting Funding fees to OFF or ON. Scanner ran the two jobs and saved the complete results. We joined the OFF and ON settings and calculated churn from the downloaded result files.
| Item | Fixed design |
|---|---|
| Markets | BTC, ETH, BNB, XRP, and SOL Binance USDT-M perpetual futures |
| Populations | Long and short separated for each market, for 10 total |
| Jobs | OFF and ON for each population, for 20 total |
| Base candle | 1 hour |
| Analysis window | 48,841 hours from 2021-01-01 00:00 UTC up to, but not including, 2026-07-29 01:00 UTC |
| Preparation data | 2,443 warmup candles outside the analysis window1, with a fixed 51,600-candle source snapshot |
| Physical upload cap | 53,001 candles including warmup |
| Entry indicators | SMA, EMA, HMA, ADX, Aroon, Vortex Indicator, Donchian Channel, Choppiness Index, RSI, CCI, Williams %R, Z-Score, ROC, CMO, Percentile Rank, R-Squared, MFI, CMF, EFI, Volume Ratio |
| Indicator periods | 5, 6, 8, 10, 12, 16, 20, 24, 32, 40, 48, 64, 80, and 96 candles for each indicator |
| Entry variants | 20 indicators × 14 periods = 280 |
| Entry rule | AND combinations of two distinct variants selected from 280, C(280, 2) = 39,060 |
| TP | 0.5%, 1%, 2%, 3%, 4% |
| SL | 0.5%, 1%, 2%, 3%, 4% |
| TTL | 6, 12, 24, 48, 96 candles |
| Exit rules | 5 × 5 × 5 = 125 |
| Results per population | 4,882,500 settings |
| Matched comparison | 48,825,000 fully identical OFF/ON setting pairs |
| Total execution | 97,650,000 backtests |
| Orders and costs | Market orders, 1× leverage, 100% capital allocation, 0.04% taker fee per fill, 0.01% slippage |
| Same-candle fill | SL takes priority if TP and SL are both touched in the same candle |
| Validation | First half of the period used in-sample, with the remaining half split into 4 out-of-sample folds |
| Result settings | Minimum trades 0, Top 100, Save all results, force a fresh run |
| Calculator version | 1.16.0 for every job |
| File checks | 20/20 result files PASS, 10/10 fully matched pair checks PASS |
Every indicator card used Scanner's Auto entry condition. The 14 periods above were the only parameters varied within each card. We did not sweep any additional indicator-specific parameters. An entry rule fired when the selected indicator-period conditions were both true. Crossing every entry combination with 5 TP values, 5 SL values, and 5 TTL values produced 39,060 × 125 = 4,882,500 settings per population.
We set the minimum trade count to 0 at execution time so the complete population would be saved. The predeclared 30-trade threshold was applied only during the primary-candidate and rank analyses. Low-trade rows therefore remained in the result files and could be filtered for the relevant analysis later.
A pair was identified by an immutable setting key built from market, direction, entry variants, TP, SL, and TTL, not by its screen rank or description. A trade-count mismatch would signal that something beyond funding had changed and invalidate the analysis. This check found no missing pairs and no trade-count mismatches.
Total return in the ON result already included the modeled funding contribution. We did not add funding income a second time. The calculation used Binance's recorded funding events and timestamps.2 It rounded millisecond timestamp jitter to the nearest minute, preserved the recorded event spacing instead of prorating a rate across hourly candles, and added each event to the first base-candle timestamp at or after that event. For each trade, the engine then took the cumulative difference over (entry, exit] and applied it to an entry-price approximation of notional value. A trade that opened and closed within the same candle received no funding.
This was not a replay of an actual account's mark-price settlement. The result should not be interpreted as an exact match for an exchange account statement.
The average barely moved, but the candidate list did
A primary candidate in this article is a setting with at least 30 trades and positive total return. We fixed this simple screening line before the analysis. The 10.8% result applies only to this boundary and does not imply the same rate at other return thresholds. It is an analytical comparison line, not certification of a live-ready strategy or future performance.
| OFF versus ON | Settings |
|---|---|
| Candidate in both | 1,338,862 |
| Candidate only with OFF | 72,154 |
| Candidate only with ON | 90,011 |
| Candidate union | 1,501,027 |
| Candidate decisions changed | 162,165 |
| Candidate churn rate | 10.8% |
About one in nine settings crossed the candidate boundary between OFF and ON. The movement did not go in only one direction. There were 90,011 settings that qualified only with ON, more than the 72,154 that qualified only with OFF. The data do not support a claim that funding simply removed candidates.
Candidate churn measures classification sensitivity at this predeclared boundary, not the economic size of every crossing. An arbitrarily small move across zero counts as one changed decision, which is why this screen identifies candidates for further validation rather than finished strategies.
A primary candidate is a configuration with at least 30 trades and positive total return. Across the ten market-direction populations, the candidate union contains 1,501,027 configurations, and 162,165 (10.8%) changed membership. 72,154 passed only with funding OFF; 90,011 passed only with funding ON. Source: candidate-transitions.csv, paired scan, 2026-08-04.
Chart 1. A primary candidate has at least 30 trades and positive total return. The two sides show the direction of movement into or out of the candidate set, not good versus bad.
An overall average total-return change of -0.0083 percentage points and a 10.8% candidate churn rate are not contradictory. When a setting sits near the zero-return boundary, a small change can determine whether it advances to the next validation stage. A setting far from the boundary can move in return while keeping the same candidate status.
The average still matters. It simply cannot reveal decision changes when the goal is to select candidates.
The Top 100 changed in some groups and barely moved in others
A moving candidate boundary did not mean the entire ranking collapsed. We compared the overlap in the highest-total-return settings between Funding OFF and ON within each population. The rank analysis included only settings with at least 30 trades. Population denominators ranged from 3,749,463 to 4,004,316 settings.
| Market and direction | Top 100 overlap | Retention |
|---|---|---|
| BTC long | 61/100 | 61% |
| BTC short | 96/100 | 96% |
| ETH long | 80/100 | 80% |
| ETH short | 79/100 | 79% |
| BNB long | 91/100 | 91% |
| BNB short | 90/100 | 90% |
| XRP long | 88/100 | 88% |
| XRP short | 96/100 | 96% |
| SOL long | 93/100 | 93% |
| SOL short | 92/100 | 92% |
BTC long retained 61% of its Top 100. The other side of the result deserves equal weight: 6 of the 10 populations retained at least 90% of their Top 100. In aggregate, 866 of 1,000 slots were retained and 134 of 1,000 changed.
Rankings include only configurations with at least 30 trades. Each row is a separate market-direction ranking; each cell reports the OFF/ON intersection and its share of K. The outline marks the 90% retention threshold. Changed slots summed across rows are not a count of distinct strategies. Source: topk-retention.csv, paired scan, 2026-08-04.
Chart 2. Each row is a separate market-direction ranking. The 134 changed slots are the combined membership changes across ten lists, not 134 distinct strategies.
Rank movement outside the very top was smaller. Among settings with at least 30 trades, the median absolute percentile-rank movement by population ranged from 0.169 to 1.284 percentage points. The p90 range was 0.561 to 3.295 percentage points. It would also be inaccurate to say that funding broadly reordered every setting that passed this filter.
Both findings matter. Most movement across the full rankings was small, while real turnover appeared at the boundary and within the leading group. Joining the two complete result sets setting by setting makes both facts visible.
Why the average canceled out: long was not always a cost and short was not always income
Breaking out average total-return differences by market and direction shows why the combined average was so small.
| Population | Average total return with Funding ON minus OFF |
|---|---|
| BTC long | -1.8747 percentage points |
| BTC short | +1.6973 percentage points |
| ETH long | -1.9761 percentage points |
| ETH short | +1.3350 percentage points |
| BNB long | +0.2751 percentage points |
| BNB short | -0.3253 percentage points |
| XRP long | -0.2965 percentage points |
| XRP short | +1.2413 percentage points |
| SOL long | -0.3750 percentage points |
| SOL short | +0.2158 percentage points |
The effect was negative for long settings in four of five markets and positive for short settings in four of five. BNB was the counterexample in both directions: positive for long and negative for short.
Funding is a signed cash flow tied to position direction, but it is not a fixed constant that can be applied across an entire study period. The sign of the reference funding rate and the intervals during which a trade was open both matter. Combining markets and directions into one average can therefore erase what happened within the individual populations.
The result does not supply a formula that subtracts a fixed cost from long positions and adds the same income to shorts. It says to apply historical funding for the relevant direction and holding intervals to the same setting, then measure the difference.
Maximum holding time and trade count were exposure proxies
Which settings were more sensitive to funding? The two proxies specified before the analysis were TTL and trade count.
TTL is the maximum holding time in candles. It is neither average nor realized holding time. Trade count is not the number of funding settlements either. Both are only indirect proxies for opportunities to be exposed to funding. For the tables below, we took the absolute value of each ON result's cumulative funding contribution, then averaged those values within each bucket.
| TTL | Primary-candidate churn | Mean absolute funding contribution |
|---|---|---|
| 6 candles | 4.723% | 3.9726 percentage points |
| 12 candles | 8.169% | 5.2343 percentage points |
| 24 candles | 10.264% | 6.1640 percentage points |
| 48 candles | 12.999% | 6.7773 percentage points |
| 96 candles | 13.961% | 7.0685 percentage points |
| Trade-count bucket | Primary-candidate churn | Mean absolute funding contribution |
|---|---|---|
| 30–99 | 2.844% | 0.3371 percentage points |
| 100–299 | 7.322% | 1.2050 percentage points |
| 300–999 | 16.533% | 4.2081 percentage points |
| 1,000+ | 27.793% | 10.2546 percentage points |
TTL is maximum holding bars, not realized holding duration. Candidate churn uses the candidate union within each bucket as its denominator; mean absolute funding contribution uses every ON result in the bucket. The two measures rising together is a conditional association, not evidence that either proxy caused the change. Source: strata-summary.csv, paired scan, 2026-08-04.
Chart 3. The denominator for candidate churn is the OFF/ON candidate union within each bucket. Mean absolute funding contribution is the simple mean across all ON result rows in that bucket. TTL and trade count are proxies for realized funding exposure, and these conditional associations are not evidence of causation.
Both candidate churn and mean absolute funding contribution rose with the two proxies in these tables. This matched the expectation recorded before analysis and did not trigger the stopping rule. Establishing longer TTL as a cause, however, would require another experiment that separates realized holding time from settlement events.
The reusable rule is simpler. The longer the maximum holding time or the more frequently a setting trades, the less willing you should be to skip an OFF/ON paired comparison.
Stronger validation made the denominator much smaller
We also tested a secondary criterion that added DSR and walk-forward conditions to the primary-candidate rule. DSR discounts Sharpe evidence for selection bias from testing many settings and for non-normality in the return distribution.3 Here it was recalculated against the complete job population within each comparison group. It is not a probability of profit.
The secondary criterion required a primary candidate to have DSR of at least 0.5 and no more than one negative out-of-sample fold.
| OFF versus ON | Settings |
|---|---|
| Passed the secondary criterion in both | 52 |
| Passed only with OFF | 89 |
| Passed only with ON | 231 |
| Union | 372 |
| Candidate churn rate | 86.0% |
The 86.0% rate looks large, but its denominator was only 372 settings. That is why it is not the headline result or part of the key takeaways. It should be read only as a caution that very few settings remained after stronger criteria were applied.
Among the fully matched pairs, a DSR difference could be calculated on both sides for 42,586,279 pairs, while 6,238,721 were missing. Walk-forward differences were valid for 43,057,497 pairs, with 5,767,503 missing. We excluded missing values from the denominator because replacing them with zero would create false observations of no change.
This small secondary result does not make the primary finding more dramatic. It shows how sharply the practical review set can narrow once validation evidence is added after candidate selection.
A paired-scan recipe for your own work
The transferable result is the comparison structure, not a particular market's number.
- Fix the signals, entry variants, TP, SL, TTL, candles, costs, and fill assumptions for one scan.
- Create two scans with the same settings and change only Funding fees from OFF to ON.
- Turn on
Save all resultsfor both before execution. This lets you download and compare the two complete result sets after the jobs finish. - Join identical settings using an immutable key built from market, direction, entry variants, and exit settings.
- Confirm that trade counts match within every pair. If they do not, stop the comparison because a rule beyond funding differs.
- Apply your candidate rule first, then split settings into both-pass, OFF-only, and ON-only groups.
- Compare candidate movement alongside Top-K retention, DSR, and walk-forward evidence. Walk-forward requires its own toggle before execution. Keep missing DSR and walk-forward values as a separate state.
Scanner's current role is to run the two configurations and save the complete results. Joining OFF and ON results on the immutable key, then calculating candidate churn and rank retention, happens in the downloaded result files.
The public audit bundle contains the analyzer, exact join key, integrity tests, input hashes, and compact output tables used here. The 3.9 GB of source Parquet pairs are not hosted in that repository, so readers can inspect the method and published aggregates but cannot independently regenerate them from the public bundle alone.
This procedure lets you create a candidate list that still qualifies after funding is included. Subtracting one estimated average cost from every setting cannot produce that list.
Scope and reproducibility limits
These results are limited to five Binance USDT-M perpetual futures markets, 1-hour candles, and the fixed analysis window. They do not establish the same magnitude or direction for another exchange, spot markets, other candle intervals, or future market conditions.
The funding calculation uses historical reference rates and an entry-price approximation of notional value. It does not replay mark-price settlement, actual account balances, or every exchange-specific settlement rule. Market impact, liquidity, and order-execution differences would also affect live fills and P&L.
Return to the opening question. Did historical funding merely nudge returns, or did it change which candidates advanced to the next validation stage?
The overall average nearly canceled out at -0.0083 percentage points. Yet 10.8% of the candidate union changed status, as did 134 of 1,000 Top 100 slots. At the same time, six of the ten populations retained at least 90% of their Top 100, and most percentile-rank movement was small.
The conclusion is not that funding is always large or always small. Pairing the same setting with Funding OFF and ON lets you build the candidate list that remains after funding is included.
Frequently asked questions
Is funding always a cost for long positions and income for shorts?
No. In this analysis the effect was negative for long settings in four of five markets and positive for short settings in four of five, but BNB reversed both patterns. The sign of the funding rate and the intervals when each trade was open both matter, so funding must be measured by market and direction.
Does 10.8% candidate churn mean the ranking collapsed?
No. Although 134 of 1,000 Top 100 slots changed, six populations retained at least 90% of their Top 100. Median absolute percentile-rank movement for settings with at least 30 trades ranged from 0.169 to 1.284 percentage points, so most movement was small.
Does a longer TTL cause funding to change more candidates?
This study found a conditional association, not causation. TTL is the maximum holding time, not realized holding time, and we used it with trade count only as a proxy for opportunities to be exposed to funding.
Does the ON result match actual exchange funding settlement?
No. It is a backtest model that applies historical reference funding rates to an entry-price approximation of notional value. It does not reproduce actual mark-price settlement, account balances, or every exchange-specific rule.
Automated trading and trading in general involve risk of principal loss. This article is for educational purposes and is not investment advice or a guarantee of returns. Historical backtest results may differ from live trading and future performance.
Footnotes
-
API automation must include the preparation window. The current public validation response does not reveal the exact candle requirement, including warmup, before a paid reservation. Follow the client calculation that adds 5% preparation candles to the analysis window, or inspect
candle_requirementsin the generated scan response. ↩ -
Binance Developers, Get Funding Rate History. This documentation identifies the source of the historical reference funding rates. It does not establish that the backtest calculation in this article matches actual account settlement. ↩
-
David H. Bailey and Marcos López de Prado, The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. The DSR in this article was calculated against the complete job population in each funding comparison group and is not a probability of profit. ↩