Correlation, regression and hypothesis testing

Reading for India · about 12 min

The answer

Correlation is 1 number that says whether 2 things moved together during a window you chose. It does not say one caused the other, it changes when the window changes, and it rises towards 1 in exactly the months when you were relying on it to be low.

Why this costs you money

A reader owns 10 holdings across several sectors, plus 2 funds. They believe they own 10 separate risks. They have done the position sizing arithmetic: 10 positions at 1% risk each, so a 10% worst case.

Then a bad month arrives and everything falls together. The loss is not 10 small independent losses. It is closer to 1 position that was 10 times too large.

Nothing was wrong with the arithmetic. It was fed the wrong number. In a calm year those holdings really did move somewhat separately, at a correlation of perhaps 0.3 or 0.4. In a fall, correlation between equities in the same market goes to 0.8 and above. The diversification was real on the days it did not matter and absent on the day it did.

The second cost is quieter. A reader tests 40 ideas, finds 1 striking result and calls it research. It was a search, and a search of that size produces striking results by chance alone.

How it works

Three tools, none of which needs notation.

Correlation

Take 2 instruments and their daily percentage changes over the same days. On a day when the first was above its own average, was the second usually above its average too?

  • +1 means always, in proportion. They are 1 thing with 2 names.
  • 0 means knowing one tells you nothing about the other.
  • −1 means the opposite, every time.

Two large Indian banks might sit at 0.7. An Indian bank and an American software company might sit at 0.2 in a calm year.

Correlation sees only straight-line relationships. Near 0 means "no straight-line relationship", never "no relationship".

Regression

Regression fits the single straight line closest to all the points. Plot your stock's daily changes against the index's, and the line gives 2 outputs.

The slope says how much the stock typically moved when the index moved 1%. A slope of 1.4 means 1.4%. That number is beta.

The R-squared says what fraction of the stock's day-to-day movement the index alone explains. An R-squared of 0.75 means 75% of the movement was market and 25% was the company. It is the more useful of the 2 and almost nobody looks at it.

The same tool on a price chart draws a regression channel: the fitted line through price, with parallel lines at the typical distance points fall from it. It carries no forecast.

Hypothesis testing

You have a claim. Say, this signal produces a higher average return than no signal.

The test starts from the opposite assumption, called the null: nothing is going on, and the difference you see is ordinary variation. Then it asks 1 question. If nothing were going on, how often would data look at least this striking?

That proportion is the p-value. A p-value of 0.03 means data this striking would appear about 3 times in 100 in a world where the effect is absent.

Read the next sentence twice.

A p-value does not tell you the probability that your idea is true. It tells you how surprising your data would be if your idea were false. Those are different questions with different answers, and swapping them is the most common error in private market research.

The 0.05 threshold is a convention, introduced by Ronald Fisher in the 1920s. There is nothing behind it.

Multiple testing, which is the real problem

Test 1 false idea at the 0.05 threshold and it passes about 5 times in 100. Test 20 false ideas and 1 of them passes on average. Not because it works, but because you asked 20 times. Test 100 and about 5 pass, and those 5 will look excellent, because the 95 you rejected are already forgotten.

Now count what you actually do. Every indicator setting you changed is a test. Every stock you checked the idea on is a test. Every date range is a test. Article 6 asks you to count them, and most active traders reach 15 to 60 over 2 years. At 40 tests, the chance that at least 1 pure coincidence clears the 0.05 line is about 87%.

The correction is simple and nobody applies it: divide the threshold by the number of tests. At 40 tests, 0.05 becomes 0.00125. One review of published stock-return predictors argued for a much higher bar for this reason.

What it tells you, and what it does not

Correlation is not causation, and there are exactly 3 alternatives. The relationship may run the other way. Both may follow a third thing you did not measure. Or it may be a coincidence produced by searching. Name which of the 4 you believe, and why.

Correlation belongs to a window. It is a property of the pair over those 250 days, not a property of the pair. A correlation quoted without a window is not a measurement.

Correlation rises in falls. This is documented. Studies of international equity markets found correlation increases in downturns and does not rise the same way in upturns. So the calm-year number is the wrong number for the purpose you computed it for.

R-squared is about the past, and small samples produce extreme results often. A fit of 0.9 describes the sample and can still be useless next month. With 20 trades almost anything can look like an edge.

The decision rule

Before you accept any statistical relationship, ask 4 questions in order.

  1. Over what window? No window stated, no measurement made.
  2. How many things did I test first? More than 10, and your threshold must

be far stricter than 0.05.

  1. What is the mechanism? If you cannot name one, treat the relationship as

a coincidence until it survives a period you did not use to find it.

  1. What is this number in the worst month? That is the only value that

matters for sizing.

Try this now

Five minutes and a spreadsheet. You are going to test your own belief that your portfolio contains different things.

  1. Pick 2 holdings you believe are genuinely different. Different sector, different size, or 1 stock and 1 fund. Choose the pair you would defend as your most separate positions.
  2. Get daily closing prices for both for the last 1 year, about 250 rows. Line them up by date in 2 columns and delete any date where one is missing.
  3. In 2 new columns compute the daily percentage change for each: today's close divided by yesterday's close, minus 1.
  4. In an empty cell type =CORREL(, select the first change column, a comma, the second change column, and close the bracket. Write the answer down.
  5. Find the worst single month of that year for the market, using the index rather than your holdings. Highlight only those 20 or so rows.
  6. Run =CORREL() again on only those rows. Write that number beside the first.

What you should see. Two numbers, and the second is much larger.

For most readers the full year lands between 0.3 and 0.6, which feels reassuring. The bad month lands between 0.7 and 0.95. If both holdings are large companies in the same market, the bad-month number is often above 0.85.

That gap is the article. You did not own 2 things. You owned 2 things on the days it did not matter and 1 thing on the day it did.

Two more minutes, and this is the part you can act on. Repeat step 4 for each holding against your home index. Any holding above 0.85 is, for risk purposes, an expensive index fund. If several sit in that group, your real number of positions is smaller than the number of lines in your holdings list, and your position sizing should use the real number.

Three real cases

1. Bangladeshi butter production and the S&P 500a perfect fit with no mechanism A researcher searched a large database of United Nations statistics for whatever best explained annual returns of the S&P 500. The winner was butter production in Bangladesh. Adding United States cheese production and the sheep population of 2 countries explained almost all the variation in the sample. The search was mechanical, the fit was genuine within the sample, and any test applied to the final model alone would have passed. 2. The Gaussian copula and mortgage securities, 2000 to 2008 — 1 correlation input priced a whole market A 2000 paper introduced a tractable way to model how likely it is that many borrowers default together. It was widely adopted for pricing pooled mortgage securities. Its correlation input came largely from a period in which American house prices had not fallen nationally. When they did fall nationally, defaults arrived together, and securities rated as very safe lost most of their value.

The model was wrong about 1 number, and that number was correlation in a stress event absent from the sample.

3. The replication project in psychology, 2015what a 0.05 threshold produces across a field A large collaboration repeated 100 published psychology experiments using the original methods. Only about 36% produced a significant result in the same direction, and average effect sizes were roughly half the originals. Those were peer-reviewed findings in a field with editors and referees. Private market research has neither.

The question that resolves it

A novice looks at 2 instruments and asks: are they correlated?

An expert asks: over what window, and what was the number in the worst 20 days?

The first has an answer that feels like a fact and behaves like an opinion. The second produces 2 numbers, and the gap between them is the risk you carry.

What would make this wrong

If correlations between equities were stable across regimes, a single historical correlation would be a sound input to position sizing and the main warning here would be unnecessary. That is an empirical claim, and the exercise above checks it on your own holdings.

Three honest limits.

Multiple-testing corrections can be too strict. Dividing the threshold by the number of tests assumes the tests are independent. Forty variations of 1 idea are not 40 unrelated ideas, so the correction will discard real effects. It is a discipline, not a precise instrument.

Correlation is not the only thing that binds a portfolio. Two holdings can show low correlation and still fail together because they share a lender, a customer or a regulator.

A p-value was designed for a fixed hypothesis and a designed experiment. Applied to price data already examined by millions of people, it means much less than the number suggests. The American Statistical Association published a formal statement in 2016 warning against this use.

In India

Correlation matters more in India because the market is more concentrated, and it is harder to measure because the data is harder to get.

Concentration. The NIFTY 50 is weighted by free-float market value, and a large share of that weight sits in financial services. An Indian portfolio built from well-known names is usually a portfolio of lenders and software exporters, whatever the individual names are. Measured correlation between such holdings is high even in calm periods.

Fund overlap. SEBI's scheme categorisation rules define mutual fund categories tightly, so many funds in a category hold similar securities. A reader who owns 3 large-cap funds from 3 fund houses has usually bought the same portfolio 3 times. Their measured correlation is often above 0.95.

The data problem. The NIFTY 50 was launched in 1996 with values computed back to a 1995 base. That is under 30 years and few complete cycles. Constituents have changed many times, and free sources rarely say what the index held on a given date, so tests silently use today's winners. Corporate actions are adjusted inconsistently across free sources. And smaller companies trade thinly, so a stale price makes measured correlation lower than the true relationship. The least liquid holdings therefore look like the best diversifiers, and they are the worst ones in a fall.

In the United States

The United States has the deepest free statistical infrastructure of any equity market, and it changes what a private person can honestly test.

Daily returns for United States stocks are available back to the 1920s through academic databases, and a widely used public library publishes daily and monthly factor returns from 1926 onward, free. Constituent history is documented and corporate actions are adjusted consistently. A reader can test on 90 years and still reserve 30 of them untouched.

Concentration exists there too and has grown. The S&P 500 is capitalisation weighted, and the largest few companies now carry a very large share of it, most in related businesses. An American investor who owns the index plus 5 well-known technology companies owns those companies twice. Correlation with the index also rose over recent decades as index funds grew, because a fund flow buys or sells every constituent at once.

Where they differ, and what that tells you

The difference is not the statistics. It is the data, and it changes what honesty requires of you.

An American reader has decades of clean, free, survivorship-corrected history. They can hold back a genuine out-of-sample period, and measure correlation across 1987, 2000, 2008 and 2020. An Indian reader has under 30 years of index history, patchy constituent history, inconsistent adjustment in free sources, and stale prices in anything below the largest companies. Reserving an out-of-sample period costs a large share of a short history, and the paid databases that fix this are priced for institutions.

What that tells you is uncomfortable. The same test carries less evidence in India than in the United States, so the Indian version of this discipline has to be stricter, not looser. Common practice is the reverse. Indian retail research usually runs on 5 years of free data from 1 website, tested on today's index members, and is quoted with more confidence than an American study built on 90 years.

So prefer a relationship with a stated mechanism over one found by searching. Treat any correlation computed on Indian small and mid-cap stocks as understated. And when you cannot reserve an out-of-sample period, write that in your notes rather than calling the backtest evidence.

Carry this

  • Correlation belongs to a window. Quote the window or you have not measured anything.
  • Correlation between your holdings rises in falls. Size on the bad-month number.
  • A p-value says how surprising the data would be if nothing were going on. It never says your idea is probably true.
  • Twenty tests at the 0.05 threshold produce 1 false winner on average. Count your tests and write the count beside the result.
  • Correlation near 0 means no straight-line relationship, not no relationship.

Knowledge check

Q. Two readers each measure a pair of holdings over the same year and each gets a correlation of 0.30.

  • Reader A holds a large private bank and a large software exporter. Both trade heavily every day.
  • Reader B holds the same bank and a small manufacturer that trades a few thousand shares a day, often with long gaps between trades.

Both conclude they are well diversified. Whose conclusion is safer?

Explanation. The first option is tempting for a sound-sounding reason. A small company in an unrelated industry really does have different customers and different costs. That reasoning is about the business. The number in front of you is about prices.

A price only updates when a trade happens. If Reader B's holding did not trade during the last hour of a falling day, its close reflects the market as it was earlier. Compare that stale close against the bank's live close and the 2 series look less connected than they are. The effect always pushes measured correlation towards 0.

So Reader B's 0.30 is partly a measurement of illiquidity, and illiquidity is not protection. On the day both want to sell, Reader A can sell. The third option fails for the same reason: the same number from 2 different mechanisms does not carry the same meaning.