What Is a Correlation and How to Measure It
SkillVeris Team
Data Science Team

Correlation measures the strength and direction of the relationship between two variables on a scale from -1 (perfect negative) through 0 (none) to +1 (perfect positive).
In this guide, you'll learn:
- Pearson correlation captures linear relationships; Spearman uses ranks and captures any monotonic relationship, making it robust to outliers and non-linearity.
- Correlation does not imply causation — a strong association can arise from coincidence or a hidden third variable.
- In Python, df.corr() builds a full correlation matrix and df['a'].corr(df['b']) measures a single pair.
- A correlation near zero rules out a linear relationship but not a curved one, so always plot the data too.
1What Is Correlation?
Correlation is a statistical measure of how strongly two variables move together, expressed as a number between -1 and +1. A value near +1 means they rise together, near -1 means one rises as the other falls, and near 0 means no linear relationship. It is one of the first tools analysts reach for when exploring how variables relate.
Correlation summarises a relationship in a single number, which makes it powerful and easy to abuse. It tells you about association and direction, but it says nothing about cause, and it can completely miss relationships that are not linear.
2Reading the Correlation Scale
The correlation coefficient always falls between -1 and +1, and both the sign and the magnitude carry meaning. The sign shows direction; the absolute value shows strength.
- +1: perfect positive — the points fall exactly on an upward line.
- 0: no linear relationship — knowing one value tells you nothing linear about the other.
- -1: perfect negative — the points fall exactly on a downward line.
- Around 0.7 to 0.9: strong; 0.4 to 0.6: moderate; below about 0.3: weak (rough, context-dependent guides).
💡Sign and Strength Are Separate
A correlation of -0.8 is stronger than +0.4. Read the sign for direction and the absolute value for strength — they answer two different questions.
3Pearson vs Spearman
The two most common correlation coefficients answer slightly different questions. Pearson measures the strength of a straight-line relationship using the actual values. Spearman first converts each variable to ranks and then measures the linear correlation of those ranks, capturing any consistently increasing or decreasing (monotonic) relationship.
- Pearson: best for linear relationships between roughly continuous, normally distributed variables.
- Spearman: rank-based, robust to outliers, and detects monotonic but curved relationships.
- Kendall's tau: another rank-based option, useful for small samples with many tied values.
When to Prefer Spearman
If your relationship looks like a curve that always trends upward, or if a few outliers are dragging Pearson around, Spearman will usually give a truer picture. Because it works on ranks, a single extreme value cannot distort it the way it distorts Pearson.
4Measuring Correlation in Python
Pandas makes correlation a one-liner. df.corr() returns a correlation matrix of every numeric column against every other, and you can pick the method with the method argument. For a single pair, use the Series .corr() method.
- df.corr() # Pearson matrix of all numeric columns
- df.corr(method='spearman') # rank-based matrix
- df['height'].corr(df['weight']) # a single pair
- df.corr()['target'].sort_values(ascending=False) # features most correlated with target
Visualising a Correlation Matrix
A heatmap makes a correlation matrix far easier to read than a grid of numbers. Libraries like seaborn plot one with sns.heatmap(df.corr(), annot=True), where warm and cool colours instantly reveal the strongest positive and negative pairs.
5Correlation Is Not Causation
A strong correlation tells you two variables move together, not that one causes the other. Ice-cream sales and drowning incidents rise together, but neither causes the other — hot weather drives both. That hidden driver is called a confounding variable.
Establishing causation requires more than correlation: a controlled experiment, a plausible mechanism, or careful causal-inference methods. Treating correlation as proof of cause is one of the most common errors in data analysis.
⚠️Beware the Confounder
Before concluding that A causes B, ask whether a third variable C could be driving both. Confounders explain a huge share of spurious correlations.
6Common Mistakes to Avoid
Correlation is easy to compute and easy to misread.
- Reading a near-zero Pearson value as 'no relationship' when the pattern is a strong curve — always plot the data.
- Applying Pearson to ranked or heavily skewed data where Spearman is more appropriate.
- Inferring causation from correlation without an experiment or mechanism.
- Ignoring outliers, which can single-handedly create or destroy a Pearson correlation.
- Comparing correlations across groups without checking each group separately (Simpson's paradox).
7Always Plot Your Data
A famous illustration called Anscombe's quartet shows four datasets with nearly identical correlation coefficients but wildly different shapes — one linear, one curved, one with an outlier. The lesson is simple: never trust a correlation number without looking at the scatter plot behind it.
- Plot a scatter for any pair that matters before quoting its correlation.
- Look for curves, clusters, and outliers the coefficient cannot capture.
- Use a pairplot to scan many relationships at once during exploration.
8Key Takeaways
Correlation is a starting point for understanding relationships, not a conclusion.
- Correlation ranges from -1 to +1: sign is direction, magnitude is strength.
- Pearson captures linear links; Spearman captures monotonic ones and resists outliers.
- In Pandas, df.corr() builds a matrix and .corr() measures a pair.
- Correlation never proves causation — watch for confounders.
- Always plot the data; identical coefficients can hide very different shapes.
9Frequently Asked Questions
Q: What is a good correlation coefficient? A: It depends entirely on the field. In physics a correlation of 0.95 might be expected, while in social science 0.4 can be meaningful. As rough guidance, absolute values around 0.7 or higher are often called strong, 0.4 to 0.6 moderate, and below 0.3 weak — but always interpret in context.
Q: What is the difference between Pearson and Spearman correlation? A: Pearson measures the strength of a linear relationship using the raw values and assumes roughly continuous, normally distributed data. Spearman converts values to ranks first, so it captures any monotonic relationship and is robust to outliers and non-linear but consistently increasing patterns.
Q: Does correlation mean causation? A: No. Correlation only shows that two variables move together. The association may be coincidental or caused by a hidden confounding variable that drives both. Proving causation requires controlled experiments, a credible mechanism, or formal causal-inference techniques.
Q: Can two variables be related but have zero correlation? A: Yes. Pearson correlation only detects linear relationships. Two variables can have a strong non-linear relationship — such as a U-shape — and still show a correlation near zero. This is exactly why you should always plot the data rather than rely on the number alone.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Data Science Team
Our data team shares real-world analytics, ML, and SQL insights grounded in industry practice.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.