Performance metrics — Sharpe, Sortino, expectancy and drawdown

Reading for India · about 13 min

The answer

Return is the least informative number about a strategy. Four numbers tell you far more: expectancy (what you make per trade on average), maximum drawdown (the worst fall from a peak), recovery time (how long it took to get back), and a risk-adjusted measure such as the Sharpe or Sortino ratio.

Of those, drawdown is the one that decides your future, because drawdown is what makes people stop. Nobody quits during a good return. They quit during a fall.

Why this costs you money

Two strategies, both returning 30% a year over 10 years.

The first fell 14% at its worst point and recovered in 5 months.

The second fell 62% at its worst point and took 3 years to recover.

On a fact sheet these are the same strategy. In life they are not remotely the same, and the difference is not comfort. It is arithmetic and it is behaviour.

The arithmetic. A fall of 14% needs a gain of 16% to get back. A fall of 62% needs a gain of 163%. Losses and gains are not symmetric, and the asymmetry gets violent as the loss grows.

Fall from peakGain needed to recover
10%11%
20%25%
33%50%
50%100%
62%163%
80%400%

The behaviour. Almost nobody sits through the second strategy. They stop at month 14 of the decline, near the bottom, and the 30% annual return in the brochure is collected by nobody. A return you cannot hold is not a return you receive.

The second, quieter cost is not knowing your own expectancy. Most traders have never computed it, so when they lose money they fix the wrong thing. They work on finding better entries when the problem is that their average loss is twice their average win. You cannot repair a number you have never measured.

How it works

Expectancy is the core number. It is what you make, on average, per trade.

Expectancy = (win rate × average win) − (loss rate × average loss)

In words: multiply how often you win by how much you win, subtract how often you lose multiplied by how much you lose.

Everything follows from this and it kills 2 popular myths at once.

A high win rate is not an edge. You can win 90% of the time and lose money, if the 10% losses are 15 times the size of the wins. This is exactly the profile of most option-selling and averaging-down approaches, and it is why they feel excellent for years and then do not.

A low win rate is not a problem. You can win 30% of the time and make a great deal of money, if the wins are 4 times the losses. This is the profile of trend-following, and it is why it is psychologically hard: you are wrong most weeks.

Two related numbers.

Payoff ratio is average win divided by average loss. Together with the win rate, it fully determines whether expectancy is positive.

Profit factor is total money made divided by total money lost. Above 1 means profitable. Below 1.2 means the result is fragile — a couple of trades going the other way would flip it. Maximum drawdown is the largest fall from a peak in your equity to a subsequent low, before a new peak. It is measured in percent and it is the single most useful risk number in existence, for a simple reason: it is the number you have to live through.

Recovery time, sometimes called time under water, is how long it took from the peak to a new peak. It gets far less attention than drawdown and it should get more. A 25% fall recovered in 4 months is an incident. A 25% fall recovered in 4 years is 4 years of your life spent making no progress while watching an index rise. That is what makes people abandon a sound strategy.

CAGR against average return. These differ and the difference is not small.

If you make 50% in year 1 and lose 50% in year 2, your average annual return is 0%. Your actual result is a loss of 25%, because 100 becomes 150 and then 75. The compound annual growth rate, or CAGR, is the constant rate that would have produced your actual final value. It is always lower than the average return when returns vary, and the gap grows with volatility. This is why volatility is a cost and not just a discomfort.

The Sharpe ratio. Your return above the risk-free rate, divided by the standard deviation of your returns. It answers: were you paid for the movement you accepted?

The measure comes from William Sharpe's 1966 study of mutual fund performance, where he called it the reward-to-variability ratio, and he revisited it in 1994. Sharpe won the Nobel Prize in Economics in 1990, though for his work on asset pricing rather than for this ratio. Cluster 9 in this wiki covers the Sharpe ratio from an investor's point of view; this article is about using it to judge a trading system.

Its limits are severe and you must know all 4.

  • It punishes upside movement. Standard deviation treats a huge gain as risk. A strategy with occasional enormous winning months is penalised for them.
  • It assumes a normal distribution. Article 4 in this cluster showed that market returns are not normal. Sharpe therefore understates the risk of anything with fat tails.
  • It can be manufactured. A strategy that sells insurance against rare events — selling far-out-of-the-money options, for example — produces small steady gains and shows a superb Sharpe ratio, right up to the event it was insuring against.
  • It depends on the sampling period. Monthly returns give a higher Sharpe than daily returns for the same strategy, because monthly sampling hides within-month movement. Always ask which was used.

The Sortino ratio fixes the first limit. It uses only downside deviation — the movement below a target return — instead of total movement. So gains are not counted as risk. The approach is associated with Frank Sortino's work in the 1980s and a 1991 paper with Robert van der Meer, and it builds on much older ideas about downside risk. For any strategy with an asymmetric return shape, which is most trading strategies, Sortino is more informative than Sharpe.

MAR and Calmar ratios divide return by maximum drawdown. CAGR divided by the worst drawdown is the most intuitive risk-adjusted number available: it tells you how much return you got per unit of pain. The Calmar ratio is usually computed over a fixed recent window, commonly 36 months. Risk of ruin is the probability that you lose enough to stop. It depends on 3 things: your edge, your position size and how much loss ends you. Of the 3, position size is the only one you fully control, and it has more effect on survival than the quality of your strategy.

What it tells you, and what it does not

These numbers describe a sample of past results. They inherit every problem from articles 3 and 6 of this cluster.

Expectancy computed over 20 trades is noise. Over 100 it is a weak estimate. It needs a large sample precisely because it is dominated by rare large results.

Maximum drawdown is a single observation. It is the worst thing that happened, not the worst thing that can happen. Shuffle the trades, as in article 6, and it changes.

None of these numbers tell you the strategy will continue to work. They describe performance, not durability.

And none of them tell you about liquidity. A metric computed on closing prices assumes you could have transacted at closing prices, in your size, on the worst day. That assumption fails exactly when the metrics matter most.

The decision rule

Compare strategies on drawdown-adjusted return, never on return.

1. Compute expectancy. If it is not clearly positive after full costs, nothing else matters.

2. Take your maximum drawdown, multiply it by 1.5, and treat that as the loss you are planning for.

3. Ask 1 question, honestly: would I still be here after that loss, for the length of time the recovery took?

4. If the answer is no, reduce position size until it is yes. Do not look for a better strategy first. Size is faster, more reliable and entirely within your control.

Then use Sortino rather than Sharpe when the return shape is uneven, and check the profit factor to see whether the whole result rests on 2 trades.

The thing being maximised is not return. It is the return you will actually be present to collect.

Try this now

This is the exercise most traders have never done, and it usually produces a surprise about which number is the problem. It needs a spreadsheet and about 5 minutes.

  1. Export your trade history from your broker's Reports, Tradebook or P&L section. Keep only closed trades. Put the net profit or loss of each into 1 column, in date order.
  2. Count the winners and the losers. Win rate = winners divided by total.
  3. Average win = the average of the positive numbers. Average loss = the average of the negative numbers, as a positive figure.
  4. Compute your expectancy: (win rate × average win) − (loss rate × average loss). This is what you make per trade on average.
  5. Compute your payoff ratio: average win divided by average loss.
  6. Now the drawdown. Add a column with the running total of profit and loss. Add another with the running total minus the highest running total reached so far. The most negative value in that column is your worst peak-to-trough loss.
  7. Find the date of that low, and the date when the running total next exceeded its previous peak. The gap between them is your recovery time.

What you should see. Three things, and the third is the one that matters.

Your expectancy is probably a smaller positive number than you expected, or a negative number, and either way it is the honest description of your trading.

Your payoff ratio is probably below 1.5, and often below 1. If it is, your problem is not your entries. You are likely cutting winners early and holding losers, which is the most common pattern in retail trade histories and is fixed in the exit rule, not the entry rule.

And your worst drawdown, with its recovery time, is a specific period of your own life. Look at those dates. Now answer the question honestly: if that had been twice as deep and twice as long, would you have continued?

If the answer is no, you have just discovered that your position size is set for a market that has already happened. That is a decision you can change this afternoon, and it does not require you to be better at anything.

Three real cases

1. Bernard Madoff (United States, collapsed December 2008)an excellent Sharpe ratio on returns that did not exist The fund reported remarkably steady positive returns with very few losing months for many years. That return shape produces an outstanding Sharpe ratio, because Sharpe rewards low variability. The steadiness was the evidence of fraud, not of skill — no genuine strategy in a liquid market produces returns that smooth. Concerns had been raised with regulators years before the collapse. The lesson is exact: a risk-adjusted metric built on reported returns measures the reporting, not the risk.

2. Long-Term Capital Management (United States, 1998)superb metrics, until the sample changed The fund produced strong returns with low measured volatility for several years, which is a high Sharpe ratio by construction. The strategy involved large borrowed positions in relationships expected to converge. When those relationships widened together in 1998, the losses were far outside anything the measured volatility implied. Low measured volatility across a calm period is not evidence of low risk. It is often evidence that the risk has not yet been sampled.

3. SEBI's studies of individual derivatives traders (India)expectancy measured across millions of accounts India's regulator has published analysis of the profit and loss of individual traders in the equity derivatives segment, reporting that a large majority lost money over the periods examined, with very substantial aggregate losses. What makes this useful here is that it is a measurement of realised expectancy, after costs, across an enormous real sample. It is the largest live test of retail trading strategies that exists anywhere, and it is published free by a regulator.

The question that resolves it

A novice compares 2 strategies and asks: which one returned more?

An expert asks: which one would I still be running after its worst year?

The first question can be answered from a chart. The second requires knowing the drawdown, the recovery time, and something about yourself. The strategy you can hold at a smaller size beats the strategy you abandon at a larger one, every time, and the gap is not close.

What would make this wrong

If drawdown did not matter, then investors would earn the returns their funds report. They generally do not. Studies of the gap between fund returns and investor returns — sometimes called the behaviour gap — consistently find that investors earn less than the funds they hold, because of when they buy and sell. That gap is drawdown, converted into behaviour, converted into money.

The honest limits are 3.

First, every metric here is computed from a sample and inherits its flaws. Expectancy over 30 trades is close to meaningless. These are tools for judging a long record, and most retail records are not long.

Second, risk-adjusted measures can mislead in the opposite direction. A strategy with a low Sharpe ratio because of large positive months is being punished for something good. Always look at the shape of the returns, not only the ratio.

Third, drawdown tolerance is personal and it is not fixed. A 40% fall in a small account you are adding to weekly is a different experience from a 40% fall in retirement savings. There is no universal acceptable number, and anybody who gives you one is selling something.

In India

Four practical points for measuring an Indian strategy.

Costs must be inside every metric. Securities transaction tax, exchange transaction charges, stamp duty, GST on brokerage and SEBI turnover fees all apply, and the treatment differs between intraday, delivery and derivatives. Expectancy computed before costs is not expectancy. For a short-term strategy, costs frequently move expectancy from positive to negative, and that is the entire result.

Taxes change the number you keep. Short-term and long-term capital gains are taxed differently, and income from derivatives is treated differently again. Two strategies with identical pre-tax expectancy can differ substantially after tax purely because of holding period.

Drawdowns can be unexitable. A stock locked at its lower circuit cannot be sold. Your computed maximum drawdown assumes you exited at the recorded prices. On the days that produce the worst drawdowns, that assumption is most likely to be false. Indian drawdown figures are therefore optimistic in a way American ones usually are not.

Lot sizes distort small-account metrics. In derivatives you cannot size below 1 lot. If 1 lot is a large share of a small account, the position size is chosen by the exchange rather than by your risk rule, and every metric you compute is a metric of that constraint.

In the United States

Standardised reporting exists. American funds report standardised performance figures, and there are established industry standards for presenting investment performance, including the treatment of composites and survivorship. This gives a retail investor reference points that are genuinely comparable, which is rarer than it sounds.

Long benchmark history. Free data allows you to compute the Sharpe ratio, drawdown and recovery time of the broad American market across many decades, so you have something to compare your strategy against. Knowing that a passive index had a drawdown of a particular depth and a recovery measured in years is essential context before you congratulate yourself on a strategy with a modest one. Cost structures are lower and simpler. Commission-free trading in shares is common, and there is no equivalent of securities transaction tax on equity trades. Slippage and the bid-ask spread remain, and for a frequently trading strategy they are still the dominant cost.

Tax treatment differs by holding period and instrument, with different regimes for short-term gains, long-term gains and certain contracts. As in India, the after-tax expectancy is the only one that pays for anything.

Where they differ, and what that tells you

The difference is that India taxes the activity and America taxes the outcome.

India's securities transaction tax and stamp duty are charged on the transaction itself, whether or not you made money. American equity trading has no comparable per-transaction tax on shares, with costs concentrated in spreads and in tax on realised gains. What that tells you is which strategies can survive in each place.

A strategy with a small edge per trade and a high number of trades can work in the United States and be arithmetically impossible in India, because the fixed cost per transaction eats an edge that is measured in fractions of a percent. It is not a question of skill. The same rule, executed perfectly, has a different expectancy in the 2 countries.

The instruction that follows is specific. An Indian trader should compute expectancy per trade and compare it directly against total cost per trade, before looking at any other metric. If your average trade makes 0.4% and your round trip costs 0.3%, you do not have a strategy. You have a job with an unusual employer. Doing this arithmetic first would prevent a large share of the losses that SEBI's studies measure.

Carry this

  • Expectancy = (win rate × average win) − (loss rate × average loss). Compute it before anything else.
  • A 50% fall needs a 100% gain. Drawdown is not symmetric and neither are you.
  • The worst drawdown you have seen is a floor, not a ceiling. Plan for 1.5 times it.
  • If you would not survive it, cut your size today. Do not look for a better strategy first.

Knowledge check

Q. Two strategies over the same 8 years.

Strategy A. Wins 72% of trades. Average win 1.4%. Average loss 3.9%. Maximum drawdown 21%. Recovery took 7 months.

Strategy B. Wins 34% of trades. Average win 6.1%. Average loss 1.8%. Maximum drawdown 27%. Recovery took 14 months.

Which has the better expectancy, and which is more likely to be abandoned by its owner?

Explanation. Do the arithmetic.

Strategy A: (0.72 × 1.4) − (0.28 × 3.9) = 1.008 − 1.092 = −0.084% per trade. A wins nearly 3 times out of 4 and loses money.

Strategy B: (0.34 × 6.1) − (0.66 × 1.8) = 2.074 − 1.188 = +0.886% per trade. B is wrong 2 times out of 3 and is roughly 10 times better than A, in the only direction that pays.

The tempting answer is the first one, and the temptation is the win rate. 72% feels like competence. It is the number people quote about themselves and it is the number that reveals the least.

Now the second half, which is where most people go wrong even after getting the first half right. B has the better expectancy and is still the one more likely to be abandoned, because B loses 2 trades out of 3, has a deeper drawdown, and took 14 months to recover. Being right about the mathematics does not make the 14 months easier.

That is the practical content of this whole article. The best strategy and the survivable strategy are different objects, and the job is not to pick between them. It is to size B small enough that you are still holding it in month 14.