Standard Deviation
Standard deviation measures how spread out the values in a dataset are from the mean, expressed in the same units as the data.
Formula
s = \sqrt{\dfrac{\sum (x_i - \bar{x})^2}{n-1}}
Definition
Standard deviation is a number that tells you how spread out data values are from the average: a small standard deviation means values cluster close to the mean, a large one means they are more spread out. The sample standard deviation $s = \sqrt{\sum (x_i-\bar{x})^2/(n-1)}$ measures average distance from the mean, with denominator $n-1$ (Bessel's correction) making it an unbiased estimator of the population standard deviation $\sigma$, expressed in the original units of measurement. Formally, $s$ is the square root of the sample variance $s^2 = \frac{1}{n-1}\sum (x_i-\bar{x})^2$; by the central limit theorem, $(\bar{x}-\mu)/(s/\sqrt{n})$ follows a t-distribution with $n-1$ degrees of freedom when the population is normal, enabling t-tests and t-intervals, and the bootstrap provides distribution-free confidence intervals for $\sigma$ when data are non-normal.
Example
Class A scores $78$, $80$, $82$ (mean $80$) have a small spread, while Class B scores $50$, $80$, $110$ (also mean $80$) have a much larger standard deviation despite the same average. For data $4$, $7$, $13$, $16$ with mean $10$: deviations are $-6$, $-3$, $3$, $6$, squared deviations $36$, $9$, $9$, $36$ sum to $90$, giving $s = \sqrt{90/3} = \sqrt{30}$, approximately $5.48$. The coefficient of variation $CV = s/\bar{x}$ expresses standard deviation as a proportion of the mean, enabling comparison of variability across datasets with different scales, such as stock volatility at different price levels.
Key Insight
Standard deviation is like the "typical distance" each value is from the mean, telling you whether most values are close together or scattered widely; the empirical rule for bell-shaped distributions says about $68\%$ of data falls within $1$ standard deviation, $95\%$ within $2$, and $99.7\%$ within $3$. Chebyshev's inequality gives a distribution-free lower bound, $P(|X-\mu| \ge k\sigma) \le 1/k^2$ for any distribution with finite variance, so at least $75\%$ of data falls within $2\sigma$ and at least $89\%$ within $3\sigma$, regardless of distribution shape.