The Strategy Score, explained
The Strategy Score is a single 0–100 number that blends a strategy's realized metrics into one at-a-glance read. It's fast triage — a way to rank a shelf of ideas in a second. It is not proof the edge will survive live. That's a different test.
What it is
The Score takes the metrics from a plain backtest and maps each one to a sub-score between 0 and 1, then takes a weighted average and scales it to 0–100. Eight components go in, each pulling the score in a clear direction:
- Return (CAGR) — weight 40, scale 0 → 18%. Deliberately the heaviest, and it never stops paying: profit beyond 18% keeps earning points instead of saturating. Beating the index is the mandate, not a box to tick.
- Drawdown — 12, scale 50% → 5% (inverted: shallower is better).
- Profit Factor — 8, scale 1.0 → 8.0. Gross profit ÷ gross loss.
- Sharpe — 8, scale 0 → 1.5. Risk-adjusted return over the risk-free rate.
- Sortino — 8, scale 0 → 2.0. Adjusted for downside risk only.
- Tail Safety — 8, scale 0.9 → 1.3. The asymmetry of the extremes.
- Stability — 8, scale 40% → 80%. The share of rolling 3-month windows that ended profitable, counting only the windows in which capital was actually at risk. A quarter spent in cash is a missing observation, not a defeat — scoring it as a loss would have made this dimension a measure of time in the market rather than of robustness across regimes.
- Exposure — 8, scale 100% → 20% (inverted). How much of the time capital was actually at work, size-weighted. This one has been argued over twice, and the argument is worth knowing: paid too generously, it rewards a strategy for inaction — trading less improves exposure and drawdown at once, two prizes for the same absence. Paid nothing, it ignores that tied-up capital is a real cost, and in a shared-capital book it is the cost that decides everything, because a sleeve sitting in cash lends its capital to the others. So it pays — but it is capped where it cannot out-vote profit. Two rules are pinned by tests: an active, profitable strategy must beat a lazy one by a clear margin, and earning more must beat de-risking without earning a euro more. Both survive at 8 and break in the mid-teens.
The weights are a profile you can change — a preservation profile leans on Drawdown and Stability, a growth profile on Return. They always sum to 100, so the number stays in 0–100.
Strategy Score = ( weighted average of 8 metric sub-scores ) × sample-confidence → [0–100]
A PORTFOLIO is scored by a different formula — a blend has no trades, so the dimensions are the ones a book turns on:
- Return 20 (0 → 18%) · Sharpe 22 (0 → 1.5) · Drawdown 22 (50% → 5%) · Alpha vs benchmark 16 (0 → 12%/yr) · Consistency 20 (50% → 85%).
- Then the whole average is multiplied by a diversification factor, so a concentrated book cannot win on raw statistics alone.
- The alpha bar is deliberately hard: 4%/yr scores 33, 8%/yr scores 67, and only 12%/yr sustained tops it out — sustained alpha of that size is extraordinary, not a starting point.
- Read the alpha with one caveat: it is measured against the benchmark without converting currencies, so a book priced in euros judged against an S&P benchmark carries some FX inside this number. Compare like with like, or read the return dimensions instead.
How to read the bands
The Score is a relative yardstick, not a precise measurement — treat it in bands, low to high. Broadly, a higher number reads Strong or Excellent, a middling one reads Decent, and a low one reads Weak or Poor. The app shows the exact scale and colours the verdict for you, so you don't have to memorize cutoffs.
The soft-curve and the ruin flag
Two design choices keep the Score honest at the extremes.
- Soft-curve. Instead of a hard cutoff, extreme metrics are compressed on a smooth curve with diminishing returns — so a Sharpe of 3.0 still edges out a 1.5, and a 90% drawdown still ranks below a 45% one, rather than everything beyond a threshold collapsing to the same value. Better metrics always help, but the marginal gain shrinks.
- Sample confidence. A great-looking score on a handful of trades is statistically fragile, so the weighted number is dampened toward zero when the trade count is small and only reaches full strength once there are enough trades. Few trades → the verdict reads inconclusive rather than shouting a grade.
- Ruin flag. A strategy that lost money on the sample, or suffered a drawdown deep enough to need a doubling just to break even, is flagged as ruinous. The flag doesn't quietly nudge the number — it's a separate gate that stops a lottery-ticket equity curve from ever displaying a top verdict.
The honest limit
This is the part that matters most. The Strategy Score reads what already happened in the backtest. It measures realized performance; it does not prove the edge will hold out of sample. A high Score on an overfit strategy is still overfit — the number just tells you the past looked good, which is exactly what an overfit curve is engineered to do.
So treat the Score as fast triage and the Overfitting Polygraph as the verdict. The Polygraph runs Walk-Forward, Monte-Carlo and a Deflated-Sharpe check and returns a real honesty read — trustworthy, fragile or curve-fit. Rank with the Score; trust only what the Polygraph clears.
Reading a Strategy Score honestly
- Is the score built on enough trades, or is the sample tiny?
- Is there a ruin flag — a net loss or a disqualifying drawdown?
- Do the component weights match what you actually care about?
- Is the high score coming from realized return, or from deep-risk metrics you'd want to survive?
- Has it passed the Walk-Forward and Monte-Carlo checks in the Polygraph?
How QUANTHEON Lab does this for you
The Strategy Score is free, computed live from each strategy's stored backtest metrics, and shown on every strategy and across the gallery — so you can compare a shelf of ideas at a glance without opening each one. Because it reads live metrics, it updates system-wide whenever the underlying numbers change. When one strategy earns a closer look, run the Overfitting Polygraph and let the verdict, not the Score, decide whether to trust it.
FAQ
What is the Strategy Score?
It's a single 0–100 number that blends a strategy's realized backtest metrics — risk-adjusted return, drawdown, tail safety, regime stability, profit factor and recovery — into one weighted average. It's a fast way to compare strategies, computed live and for free from stored metrics.
Does a high Strategy Score mean the strategy is good?
No. The Score only reads what already happened in the backtest; a high Score on an overfit strategy is still overfit. Use it as triage, then confirm the edge with the Overfitting Polygraph — Walk-Forward, Monte-Carlo and a Deflated-Sharpe check — which returns the real trustworthy / fragile / curve-fit verdict.
Is the Strategy Score free?
Yes. It's computed live from a strategy's stored base-backtest metrics and shown on every strategy and across the gallery at no cost — it doesn't depend on any paid feature.
Related: What is overfitting? · How to backtest a strategy · Sharpe ratio calculator · Max drawdown calculator