Variance in Statistics

What variance measures, how to calculate sample and population variance step by step, why its units are squared, and when to report variance instead of SD.

Variance in Statistics

You need to know how spread out a set of numbers is, but you need the number squared, and you need to know whether you are looking at a sample or the whole population. That is what variance is for. The variance is the average of the squared deviations from the mean. It is the square of the standard deviation, and it is the number you calculate before you take the square root to get the standard deviation. This covers the variance formula, how to calculate variance, the difference between sample variance and population variance, and how variance compares to standard deviation.

OpenStax Introductory Statistics 2e (section 2.7) defines the standard deviation as a number that measures how far data values are from their mean, and the variance as the square of the standard deviation. The sample variance uses the symbol s², and the population variance uses σ². The formula for sample variance is s² = Σ(x - x̄)² / (n - 1), where n-1 is called the degrees of freedom. The formula for population variance is σ² = Σ(x - μ)² / N. The denominator is the key difference: n-1 (Bessel's correction) for a sample, N for the population. Using n-1 makes s² an unbiased estimator of σ², though the sample standard deviation s is not an unbiased estimator of σ.

Variance Formula: Population Vs Sample

The variance formula is the same idea in both cases: it is the sum of squared deviations from the mean, divided by the number of data points. For a population, you divide by N. For a sample, you divide by n-1. The n-1 is there because you have already used the sample mean to calculate the deviations, which uses up one degree of freedom. You are left with n-1 independent pieces of information. This correction means that s², on average, equals σ². If you used n in the sample formula, you would get a biased estimate that is too small on average.

The most common mistake newcomers make is using the population formula (divide by N) on a sample. That underestimates the true population variance. Using the sample formula (divide by n-1) on a full population overestimates it. Check which you have before you type: if your data is the entire group you care about, use N. If it is a sample meant to estimate a larger group, use n-1. The population variance formula is σ² = Σ(x - μ)² / N. The sample variance formula is s² = Σ(x - x̄)² / (n - 1).

Worked Example: Calculating Variance By Hand

Take this data set: 4, 8, 6, 5, 3. You will treat it as a sample. First, find the sample mean x̄. The sum is 26, and n is 5, so x̄ = 26 / 5 = 5.2. Now find the deviations from the mean: 4 - 5.2 = -1.2; 8 - 5.2 = 2.8; 6 - 5.2 = 0.8; 5 - 5.2 = -0.2; 3 - 5.2 = -2.2. Square each deviation: (-1.2)² = 1.44; 2.8² = 7.84; 0.8² = 0.64; (-0.2)² = 0.04; (-2.2)² = 4.84. Sum the squared deviations: 1.44 + 7.84 + 0.64 + 0.04 + 4.84 = 14.80. Divide by n-1, which is 4: 14.80 / 4 = 3.70. The sample variance s² is 3.70. The units are squared, so if the original data was in dollars, the variance is in dollars squared.

If this were a population, you would divide by N = 5: 14.80 / 5 = 2.96. The population variance σ² would be 2.96. Notice the population variance is smaller than the sample variance. That is because dividing by a larger number (N) gives a smaller result. The sample variance's denominator (n-1) is smaller, so the estimate is larger, which corrects for the bias introduced by using the sample mean instead of the population mean.

Variance Vs Standard Deviation: Why Both Exist

Variance and standard deviation measure the same thing: the spread of data around the mean. The standard deviation is the square root of the variance. The variance is in squared units, so if your data is in metres, the variance is in square metres. That is hard to relate to the original data. The standard deviation is in the original units, so it is easier to interpret. For the worked example above, the sample variance was 3.70. The sample standard deviation s is √3.70 ≈ 1.92, which is in the same units as the original data.

Keep Variance For Math, Report Standard Deviation

The variance is the number you need for further statistical calculations. Analysis of variance (ANOVA) works on variances, not standard deviations. When you add independent variables, the variances add, not the standard deviations. The standard deviation is the number you report for interpretability because it is on the same scale as the data. When someone asks how spread out the data is, you give the standard deviation. When you are doing the math that leads to that number, you work with the variance. The confusion is normal at the start. Remember: you calculate the variance first, then take the square root to get the standard deviation.

Where Variance Is Used: ANOVA And Adding Independent Variables

Variance is the building block of many inferential statistics procedures. Analysis of variance (ANOVA) compares the means of three or more groups by looking at the ratio of between-group variance to within-group variance. If the between-group variance is large relative to the within-group variance, the group means are likely different. ANOVA works because variances add under certain conditions, and that property does not hold for standard deviations.

Additivity Of Variances

When you add two independent random variables, the variance of the sum is the sum of the variances. For example, if you have one variable with variance 4 and another independent variable with variance 9, the variance of their sum is 4 + 9 = 13. The standard deviations would be 2 and 3, but the standard deviation of the sum is √13 ≈ 3.61, not 2 + 3 = 5. This additivity of variances is the mathematical reason variance is used in the formulas for error propagation, portfolio risk in finance, and the calculation of the standard error of the mean (SEM = σ / √n).

The standard error of the mean is the standard deviation of sample means. To get it, you divide the population standard deviation by the square root of the sample size. The formula works because the variance of the sample mean is σ² / n, and then you take the square root to get the SEM. If you tried to work with standard deviations directly, the math would break down. This is why variance is a fundamental quantity in statistics, even though most people report standard deviations.

Why The Units Are Squared And What That Means

The variance formula squares each deviation because a simple average of deviations from the mean is always zero. The positive deviations cancel the negative ones. Squaring makes every deviation positive before averaging. The side effect is that the variance is in squared units. If you are measuring test scores out of 100, the variance is in points squared. A variance of 100 points squared does not directly tell you how far a typical score is from the mean. You take the square root to get the standard deviation of 10 points, which you can interpret.

This squared-unit property is not a flaw. It is what makes variance additive. When you square a deviation, you lose the intuitive scale, but you gain mathematical convenience. The trade-off is that you always report the standard deviation for interpretability and keep the variance for computation. The coefficient of variation (CV = σ / μ × 100%) is a unitless measure of relative variability that uses the standard deviation, not the variance, because the units cancel when you divide by the mean.

Common Questions

Why is the variance in squared units?

Because you square the deviations before averaging them. A deviation is in the original units, so its square is in units squared. This makes the variance difficult to interpret directly, which is why the standard deviation (the square root) is often reported instead.

What is the difference between sample variance and population variance?

Sample variance (s²) uses n-1 in the denominator to correct for bias when estimating a population parameter from a sample. Population variance (σ²) uses N because the data is the entire group, so no correction is needed.

How do I know which variance formula to use?

If your data set includes every member of the group you want to describe, use the population formula (divide by N). If your data is a sample selected to represent a larger group, use the sample formula (divide by n-1).

Is the sample variance an unbiased estimator?

Yes, s² is an unbiased estimator of σ². However, the sample standard deviation s is not an unbiased estimator of the population standard deviation σ. The bias in s is small for large sample sizes (n > 30) but can be several percent for small samples.