Statistics for technicians

Reading for India · about 12 min

The answer

Almost every indicator you have met is 1 of 3 statistics wearing a chart. The mean is the middle. The standard deviation is the usual distance from the middle. Correlation is whether 2 things move together.

Learning those 3 properly is worth more than learning 20 indicators, because every indicator is a rearrangement of them — and because the place they fail is the same place your indicators fail.

Why this costs you money

You size a position using a rule that feels sensible. You know your stock "usually moves about 2% a day", so you set a stop 4% away and you assume the worst case is somewhere near 6%.

Then a day arrives where it falls 19%.

Your reaction is that this was extraordinary, a once-in-a-lifetime event, bad luck. It was not. Markets produce these days at a rate that is many times higher than the standard statistical model predicts, and the model you were using without knowing it is the model that says they are nearly impossible.

This is not an abstract complaint. It is the mechanism behind most account destruction. Position sizing, option pricing, stop placement, risk limits, value at risk models used by large institutions — all of them, in their simplest form, assume that extreme moves are rarer than they are. The person who understands this sizes smaller and survives. The person who does not is correct 200 times and finished once.

The second cost is smaller and constant. Not knowing what "2 standard deviations" means turns Bollinger Bands into a mystery ritual, when they are just a moving average with a moving ruler attached.

How it works

The mean. Add the values, divide by how many there are. On a price chart, the mean of the last 20 closes is a 20-day moving average. That is the entire idea. A moving average is not a prediction and it is not support. It is the middle of a window, recalculated as the window slides.

The standard deviation. This is the number that unlocks everything, so here it is without notation.

Take a set of daily moves. Find the average move. For each day, work out how far it was from the average. Some are above and some below, so square each of those distances to make them all positive. Average the squares. Then take the square root to get back to the original units.

What you end up with is the typical distance from the middle. If the average daily move of a stock is 0 and its standard deviation is 1.4%, then a normal day is a move of roughly 1.4% in either direction.

In markets, standard deviation of returns has a name: volatility. They are the same thing.

Two practical facts about it.

It scales with the square root of time, not with time. If daily volatility is 1%, then 4-day volatility is not 4%. It is 2%, because the square root of 4 is 2. To convert daily volatility to annual, multiply by the square root of the number of trading days in a year, roughly 252, which is about 15.9. This is why an annualised volatility of 16% and a daily move of about 1% are the same statement.

It is symmetric, and markets are not. Standard deviation treats a 5% rise and a 5% fall as equally significant. Your account does not, and article 7 in this cluster explains what to use instead.

The normal distribution. This is the familiar bell shape. It has a convenient property: if data follows it, then about 68% of observations fall within 1 standard deviation of the mean, about 95% within 2, and about 99.7% within 3.

That last number is the one to remember. Under a normal distribution, a move larger than 3 standard deviations should happen about 3 days in every 1,000. A move larger than 5 standard deviations should happen roughly once in several thousand years.

Market returns are not normal. They have what is called fat tails: the extreme outcomes are far more common than the bell shape predicts. Small moves cluster more tightly than the model says, and huge moves happen far more often than it says. Both halves of that sentence are true and the second half is the one that costs money.

A worked illustration you can check. On 19 October 1987 the Dow Jones Industrial Average fell about 22.6% in 1 session. Measured against the volatility of the preceding period, that is somewhere above 20 standard deviations. Under a normal distribution, an event of that size would not be expected within the lifetime of the universe. It happened on a Monday.

Volatility clusters. Large moves are followed by large moves, and quiet periods are followed by quiet periods. This was described as early as the 1960s and is the basis of a whole family of statistical models. The practical consequence is direct: today's volatility is the best simple forecast of tomorrow's, and article 9 turns that into a position-sizing rule.

Bollinger Bands, decoded. A middle line, which is a 20-period moving average — the mean. An upper and lower band placed 2 standard deviations away — the ruler. That is all. Under a normal distribution about 95% of observations would fall inside those bands. In practice, on real price data, the observed proportion is lower, around 88 to 89%. The gap between 95% and 89% is the fat tail, visible on your own screen. It is also why "price touched the upper band" is not a sell signal: touching the band is a normal event that happens constantly, not a rare one.

What it tells you, and what it does not

These statistics describe a sample. They do not describe the world.

Every number you compute is from a window you chose. A 20-day volatility and a 200-day volatility of the same stock can differ by a factor of 3, and neither is wrong. The window is an assumption, and it is usually an invisible one.

They tell you nothing about direction. Volatility is a measure of size, not of sign. A stock with rising volatility is not falling. It is moving more.

They assume the past sample contains the range of possibilities. It does not. Your data covers whatever happened to occur, and the largest event in your sample is not the largest possible event. It is the largest one so far.

And an average can describe a situation that never occurred. If a stock rose 30% in half its months and fell 25% in the other half, its average month is positive and there was no such month.

The decision rule

Convert every risk decision into standard deviations before you make it.

1. Find the daily standard deviation of the instrument, over a window you state out loud.

2. Express your stop distance as a multiple of it. A stop 1 standard deviation away will be hit by noise, constantly. A stop 3 away will rarely be hit by noise, and costs more when it is.

3. Express your position size so that a 5 standard deviation day — which the normal model says is impossible and reality says happens — is survivable.

Then say, in 1 sentence, what window you used. If you cannot, you do not know what your numbers mean.

The reason to work in standard deviations rather than percentages is that it makes different instruments comparable. A 3% stop is loose on a quiet large company and extremely tight on a volatile small one. "1.5 standard deviations" means the same thing on both.

Try this now

You are going to measure your own stock and then find the days that should not exist. This needs a spreadsheet and about 5 minutes.

  1. Pick 1 stock or fund you own. Download or copy its daily closing prices for the last 6 months. Most broker platforms and free chart sites export this. About 120 rows.
  2. In the next column, compute the daily percentage change: today's close divided by yesterday's close, minus 1, times 100.
  3. At the bottom, compute the average of that column and its standard deviation. In a spreadsheet these are =AVERAGE(range) and =STDEV(range).
  4. Multiply the standard deviation by 3. Write that number down.
  5. Now count how many days in your 120 had a move larger than that number, in either direction.

What you should see. Under the normal distribution, 120 days should contain 0 such days — about 3 in every 1,000, so you would expect to find one about once every 18 months. Many readers will find 1. Some will find 3 or 4.

Every one of those is a day that a standard risk model said would essentially never happen, on a stock you own, inside 6 months.

Second, quicker check. Multiply your daily standard deviation by 15.9. That is the annualised volatility of your holding, and it is the number that would appear in a fund fact sheet or an option pricing screen. Compare it between 2 of your holdings. The larger number is the one that will test your patience, regardless of which has the better business.

Three real cases

1. Anscombe's quartet, 1973identical statistics, 4 different worlds

A statistician published 4 small datasets that share almost every summary number: the same mean for both variables, nearly the same standard deviation, nearly the same correlation, and the same fitted straight line. Plotted, they look nothing alike. One is a clean straight relationship. One is a curve. One is a straight line ruined by a single outlier. One is a vertical cluster plus 1 distant point that creates the entire apparent relationship. It remains the best 30-second argument ever made for looking at the data before trusting a statistic computed from it.

2. 19 October 1987 (United States)the impossible day The Dow Jones Industrial Average fell approximately 22.6% in 1 session. Measured in standard deviations of the preceding period, the size of that move places it far outside anything a normal distribution admits. No new information of corresponding size arrived. The relevance is not the history. It is that every risk model in use that morning treated the day as impossible, and the model was not repaired by the day being called an exception.

3. Long-Term Capital Management, 1998 (United States)statistics applied by people who understood them The fund's positions were built on measured relationships with strong statistical support. In 1998 many of those relationships moved adversely at the same time. Correlations that had been low in the sample went to nearly 1 in the event, which is a statistical statement about a real disaster. The people involved were not naive about statistics. They were among the best in the world at it, and that is the point worth carrying: understanding the mathematics does not protect you from the assumption underneath it.

The question that resolves it

A novice sees a large move and asks: why did it happen?

An expert asks: how large was it, in standard deviations, on what window?

The first question produces a story, and there is always a story available. The second produces a number that can be compared with other moves, other instruments and other periods. Only 1 of those 2 answers is still useful next month.

What would make this wrong

If market returns were normally distributed, then over a long enough history the number of 3 standard deviation days would match the prediction of about 3 per 1,000. Count them on any index with 30 years of data. The observed count is consistently far higher.

The honest limits are 3.

First, "fat tails" is a description, not a model. Knowing that extremes are more common than normal does not tell you how much more common. Several competing statistical models exist and they disagree. Do not replace false precision with a different false precision.

Second, standard deviation is still useful. This article is not an argument for abandoning it. It is the best simple summary of typical movement available, and it is what everything from option pricing to position sizing runs on. The correction is to treat it as a description of ordinary days and to size for extraordinary ones separately.

Third, the window changes the answer. Anybody, including me, can produce the volatility number they want by choosing the period. Always state it.

In India

Indian markets have 3 features that change these statistics in ways worth knowing.

Circuit limits truncate the tails, on paper. Individual stocks have price bands, commonly 5%, 10% or 20%, and the index has market-wide circuit breakers at 10%, 15% and 20%. A stock that would have fallen 40% falls 20% and stops. Your data then shows a 20% day. The recorded standard deviation is understated, because the true move was not allowed to print. A backtest run on that data assumes you could have exited at the closing price. You could not — a stock at its lower circuit has no buyers.

Volatility is generally higher than in developed markets. The NIFTY 50 has historically moved more than the S&P 500 in a typical year. This matters when you copy a rule from an American source: a stop distance calibrated for American volatility is a tight stop in India.

Expiry sessions distort short-window statistics. Because a very large share of Indian volume concentrates into index option expiry, intraday volatility on those days is not drawn from the same distribution as other days. If you compute a 20-day statistic that includes 3 expiry sessions, you have mixed 2 different populations. For intraday work this is a real problem and it is usually ignored.

Free data for this work is available: NSE publishes daily price files and historical index values, and most broker platforms allow an export to a spreadsheet.

In the United States

American markets give you the longest and cleanest series in the world for this kind of measurement, and several are free.

Broad market return series reach back to 1926 in widely used academic datasets, and 1 well-known dataset of prices and dividends reaches to 1871. That length lets you do something Indian data mostly cannot support: measure how the statistics themselves change across eras.

Three American-specific points.

No general daily price limits on individual shares. American shares can fall as far as buyers allow in a session, so the recorded tail is the real tail. There are trading pauses under the limit-up limit-down mechanism and market-wide circuit breakers, but the routine truncation that Indian price bands create does not apply in the same way. American data is therefore a more honest record of extreme moves.

Volatility regimes are well documented. Long stretches of very low volatility, such as much of the mid-2010s, and short violent ones, such as February 2018 and March 2020, are all in the free record. Studying the transitions is one of the few genuinely useful exercises available to a retail researcher.

The measured statistics are widely published. Realised volatility, implied volatility and their difference are computed and published by exchanges and data vendors, so you can check your own arithmetic against a reference.

Where they differ, and what that tells you

The difference is that India truncates its tails and America records them.

Price bands and frequent circuit halts mean the Indian historical record systematically understates how far prices actually wanted to move on the worst days. American data mostly does not.

What that tells you is where each market's risk is hidden.

In the United States, the danger is visible in the data and people underestimate it anyway, because the normal distribution is comforting. The fix is statistical humility.

In India, the danger is partly absent from the data. A backtest of an Indian strategy through a crisis will show losses that were survivable, because the day the position could not be exited at any price appears in the file as a single 20% bar. The risk you cannot see in the data is liquidity risk, and Indian statistics hide it more than American ones do.

The practical conclusion is the same in both places and it is why this article sits early in the cluster: size positions so that the worst day in your data is not the worst day you can survive.

Carry this

  • Standard deviation is just the typical distance from the middle.
  • Volatility scales with the square root of time. Daily times 15.9 is annual.
  • Real markets produce 3 standard deviation days far more often than the model says. Count them on your own holding.
  • Every statistic is computed on a window. State the window or you do not know what the number means.

Knowledge check

Q. Two stocks, over the last 6 months.

Stock A has a daily standard deviation of 1.2%. Its largest single-day move in the period was 3.8%.

Stock B has a daily standard deviation of 1.2%. Its largest single-day move in the period was 14%.

Both have the same volatility number. Which one is riskier for a position sized by that number, and why?

Explanation. The tempting answer is the third one. Standard deviation is computed from every day including the extreme ones, so it feels as though the 14% day is already inside the number. It is, arithmetically. It is also almost invisible in it, because 1 large day in 120 barely moves the average.

The 2 stocks have the same typical day and completely different worst days. A stop placed 3 standard deviations away, at 3.6%, would have held on A and would have been jumped entirely on B — the price would have opened past it, and the loss taken would be far larger than the loss planned.

That is the whole practical content of fat tails. Volatility tells you about ordinary days. It does not tell you what happens on the day your stop does not work. Sizing from volatility alone assumes those 2 things are related, and in B they are not.

The first option is the gambler's error in respectable clothing. A stock is never "overdue". The last 120 days have no memory and no obligation.