OptiNod Academy

The Three Robustness Scores — How to Decide Which of Two Equal-Return Settings to Trust

How OptiNod's three robustness scores (Neighborhood Stability, Edge Consistency, Top-N Consistency) are calculated, and how to read the overall score and parameter sensitivity when choosing between two settings with the same return.

Whether the top return can be trusted is decided by the results of the settings around it and by how all the tested combinations are distributed.


Robustness originally comes from statistics, where it describes an estimate that barely moves when an assumption does not quite hold. In trading it means that performance stays roughly the same when a parameter is moved by one step or the data period is changed a little. OptiNod's robustness analysis scores the whole set of optimization results in three ways and combines them into a single overall score from 0 to 100.


Many traders treat this score as a formality to glance at after the backtest. They sort the results table by return, pick the top row, and use it as long as the robustness score shows no warning. The robustness score, however, does not rate the top row on its own. It rates whether this optimization run contains a region of settings that can be trusted. When the score is low, this run has no setting worth choosing, however high the top return is. The overfitting article explains why a strong top result with weak neighbors appears.


This article explains what each of the three scores calculates, how to read the overall score and the parameter sensitivity table, and how to apply the same logic when choosing between two settings with the same return. The aim is to check the robustness score and its three components before looking at the return column.


Spike-shaped versus plateau-shaped performance
Spike-shaped versus plateau-shaped performanceEven with the same top return, performance that drops sharply around the best parameters is likely chance, while performance that holds across a wide region is more likely to hold in live trading.

Even performance around the top settings raises Neighborhood Stability


The first score, Neighborhood Stability, looks at how similar the results around each of the 10 best results are. A result counts as nearby when it is close on every numeric parameter. The distance limit is the larger of the parameter's step and 10% of the tested range, and select or checkbox parameters must have the same value. The calculation runs only when a neighborhood holds at least three results, including the top result itself. The standard deviation of the neighborhood's performance is compared with the spread of all results, and a smaller local spread earns a higher score. The overall spread is measured with the interquartile range, so a few extreme results do not dominate it. The top result, which is the setting a user applies, makes up half of the score, and the average of the other top results makes up the other half. This score carries 35% of the overall score.


The score matters because performance that reflects a recurring market behavior stays close when a parameter moves by one step. Changing an EMA period from 20 to 21 does not make a trend disappear. When the settings next to the top one drop sharply, the top result most likely matched a few specific candles by chance. Those candles will not return in live trading, so the setting's performance soon disappears.


On May 19, 2021, BTC opened at $42,850, fell to $30,000 intraday, and closed at $36,690. A setting whose stop-loss distance is tuned to the tick for a single crash candle like this easily becomes the top backtest result. Widening the stop by one step, however, lowers its performance sharply, and Neighborhood Stability picks up that difference. Look at more than the result of one stop level: check how much the result changes when the stop is moved one step in each direction.


Edge Consistency is low when most combinations lose money


The second score, Edge Consistency, looks at how many of all tested combinations beat the break-even level. The break-even level depends on the metric: 0 for net profit and the Sharpe ratio, and 1 for the profit factor. Sixty percent of the score comes from the share of combinations above that level, which earns nothing at 30% or below and full marks at 90% or above. The remaining 40% comes from how far the median result of all combinations sits above break-even. This score also carries 35% of the overall score. When the analysis uses a metric without a break-even level, such as maximum drawdown or win rate, this score is skipped and its weight is shared between the other two.


The question behind this test is simple: is good performance spread across the parameter space, or does it exist only in a few selected combinations? If, on net profit, only 35% of the combinations made money and the median is a loss, the top return was picked from a set of mostly losing combinations. Earlier versions of OptiNod used a Monte Carlo test in this place. It compared the top results with random subsets drawn from the same results, and it returned a p-value near zero even when every combination lost money, so it could not support a decision. The current test counts directly how many combinations beat break-even.


This test shows when the strategy itself does not make money in the tested period. During the FTX collapse in November 2022, BTC fell from $18,545 to an intraday low of $15,588 on November 9. Optimizing a long strategy on bear-market data that includes a crash like this leaves most parameter sets losing money, and Edge Consistency comes out low. That is not a parameter-selection problem. It signals that this strategy does not fit this period. Checking this share first is the way to narrow the gap between backtests and live trading.


Where the top results sit: clustered means a real peak, scattered means chance
Where the top results sit: clustered means a real peak, scattered means chanceThe top-performing settings are plotted on a parameter map. When they gather in one area, that area is where performance is genuinely good; when they are spread across the map, the top result is closer to chance.

Top results gathered in one place raise Top-N Consistency


The third score, Top-N Consistency, looks at whether the parameters of the top 20% of results (at least three when there are few results) sit in a narrow range. Sixty percent of the score is concentration. For each parameter, the range covered by the top results is compared with the full tested range, and a narrower range earns a higher score. Parameters that move performance more are weighted more heavily, because a parameter that barely changes performance can be spread widely among the top results without being a sign of overfitting. The remaining 40% is the share of top results in the largest connected group, using the same distance limit as Neighborhood Stability. This score carries 30% of the overall score.


When the top results gather in one area, the market has a region of settings that genuinely performs. If the top EMA periods all fall between 18 and 24, that range is where performance is good. If the top results are spread across periods 10, 50, and 90, each result is more likely a separate piece of chance, and none of them is guaranteed to repeat in the next period.


Take March 14, 2024, when BTC set an all-time high of $73,777 and then fell to $68,555 on the same day, a move of more than $5,000. In a period like this, settings that happened to catch one or two large candles reach the top while staying scattered. A low share for the largest group means the top results did not form a single region and several pieces of chance are mixed together. Splitting the data with a walk-forward analysis shows these scattered top results moving to different places in each segment.


The overall score rates the whole run; parameter sensitivity rates each parameter


The three scores are combined with weights of 35%, 35%, and 30% into the overall score. No score is calculated when there are fewer than five distinct results or when every result has the same parameter values, because judging stability from too few samples is itself another form of overfitting. The overall score shows the following messages:


Overall scoreMessage
80 or higherResults appear robust. Low risk of overfitting.
60–79Results are moderately robust. Consider forward testing to validate.
40–59Results show some signs of overfitting. Proceed with caution.
Below 40Results may be overfit. Strongly recommend reducing parameters or forward testing.

The Parameter Sensitivity table scores each parameter separately. A parameter scores lower the more it moves performance and the more widely its values are spread among the top results. Scores below 60 are marked "Moderately sensitive" and scores below 40 "Highly sensitive"; for a highly sensitive parameter the table suggests widening its range and testing again, or fixing its value.


Choosing between two settings with the same return


Because the robustness score rates the whole run, comparing two candidates means applying the same logic to each candidate directly. The thresholds below are recommendations; the product does not enforce them. Adjust them to the character of the strategy.


  1. Check the overall score: Below 40, choose neither candidate. The run has no region worth trusting, so change the strategy or the test period first.
  2. Check Edge Consistency: If fewer than half of the combinations beat break-even, parameter selection will not solve the problem. Review the strategy's trading rules first.
  3. Check each candidate's neighbors: In the heatmap or the results table, look at the results with each candidate's inputs moved one step. Keep the candidate whose neighbors have a higher average and a smaller spread.
  4. Deal with sensitive parameters: Fix the value of any parameter marked "Highly sensitive" in the Parameter Sensitivity table, or widen its range and run again.
  5. Final comparison: Among the remaining candidates, choose the one with the higher profit factor and the lower maximum drawdown, then run it once more on a period the optimization did not use.

Two common ways to misread the scores


Looking only at the overall score. An overall score of 60 can come from three scores in the 60s, or from one score of 90 and two in the 40s. The second case passed only one kind of check. Use the overall score as a starting point, and do not skip opening the three components to see which check is weak.


Trusting the score when there are few results. The score is calculated with as few as five results, but with few results fewer top results have the three neighbors needed for Neighborhood Stability. A high score in that case rests on weak evidence. Before trusting the score, run enough combinations to have at least 30 results.


Trust the results only when all three scores are high


A robustness score does not promise future returns. It separates good performance that appears across a wide region of the backtest from performance that appears only in a few combinations that matched by chance. The final check is therefore whether all three scores are high. The thresholds below are also recommendations.


  • Is the overall score 60 or higher?
  • Is none of Neighborhood Stability, Edge Consistency, and Top-N Consistency below 40?
  • Did at least half of the combinations beat break-even in Edge Consistency?
  • Is no parameter marked "Highly sensitive" in the Parameter Sensitivity table?
  • Are there at least 30 results, so the scores have enough resolution?

Results with all three scores high are more likely to hold similar performance when the data period changes. They are different from a setting fitted to a single shock candle, such as BTC falling from $58,161 to $49,000 on August 5, 2024, when the yen carry trade unwound. Which of two equal-return settings to trust is decided by these three scores together with the results around each candidate.

Check your own backtest

Upload an optimization results CSV or a TradingView strategy trade list to see its performance metrics together with robustness checks. No sign-in is needed to analyze.

Analyze my results file Explore an example report