Knowing the centre of a dataset tells only half the story; two datasets can share an identical mean yet behave completely differently if one is tightly clustered and the other wildly scattered. Measures of spread exist to quantify this scatter — how far observations typically stray from the centre — and without them, analysts cannot distinguish a reliable, consistent process from an erratic, unpredictable one. Before spread was formalised, comparisons rested entirely on averages, which hid risk and volatility entirely. Variance, standard deviation, and the interquartile range each capture dispersion in a different way, and they underpin everything from risk modelling and quality control to the confidence intervals and hypothesis tests that form the backbone of inferential statistics.
25 minintermediate
Measures of Spread — Variance, Std Dev, IQR
Analogy🏏Cricket
🏏 Think of it like cricket: Imagine Virat Kohli has scored 45, 52, 38, 61, and 49 across five innings in a series, and a commentator wants to describe his form in one phrase. The commentator cannot read out all five scores every time, so they compress them into a single representative figure. Just as the commentator picks one number to stand in for the whole sequence of innings, a measure of central tendency picks one value to represent an entire dataset. Just as a misleading summary ('he averages 200') would distort how selectors judge Kohli, a wrongly chosen centre distorts how analysts judge data. And just as different summaries (best score versus typical score) tell different stories, the mean, median, and mode each emphasise a different aspect of the same innings. This reveals why central tendency is never one fixed number — it is a deliberate choice about which story the data should tell.
Lesson 2 of 35
0% complete