Mean, Median and Standard Deviation for Grouped Data
Calculate the mean, median and standard deviation from a frequency table or class intervals using midpoints, with a full worked example and common slips.
You Have a Frequency Table, Not the Raw Data
A student posts a problem: "Here is a frequency table of exam scores. Find the mean, median, and standard deviation of grouped data." The raw scores are gone. You have a list of class intervals and a count (f) for each. The formulas and the failure modes when you can only approximate are explained.
Frequency Tables vs. Class Intervals
Data are grouped into classes when the original values are too many or too detailed to list. A frequency table assigns a frequency, f, to each class. If the data are discrete, each row might be a single value. More often, data are continuous and the table uses class intervals, such as 60-69.
Each class has a lower class limit and an upper class limit. The class width is the gap between the lower limit of one class and the lower limit of the next. The class midpoint, xm, is (lower limit + upper limit) / 2. This midpoint stands in for every value in that class, which introduces estimation error. The finer the class width, the better the estimate.
OpenStax Introductory Statistics 2e, sections 2.6-2.7, uses frequency tables to introduce these summaries. The same sections define a frequency table as a listing of each data value (or group of values) and its frequency.
Mean of Grouped Data
To estimate the mean of grouped data, treat each class midpoint as the representative of all values in that class. Multiply each midpoint by its frequency, sum those products, then divide by the total number of data points, n.
Formula: x̄ = (Σ xm * f) / n
Work through an example. Fifty students took a test. The frequency table is:
- 60-69: f = 5, xm = 64.5
- 70-79: f = 10, xm = 74.5
- 80-89: f = 20, xm = 84.5
- 90-99: f = 15, xm = 94.5
Compute Σ xm * f = (64.5 × 5) + (74.5 × 10) + (84.5 × 20) + (94.5 × 15) = 322.5 + 745 + 1690 + 1417.5 = 4175. n = 50. The estimated mean of grouped data is 4175 / 50 = 83.5.
Check Against the Calculator
Most calculators that do 1-Var Stats can accept frequency lists. Enter the midpoints as the data list and the frequencies as the frequency list. The calculator returns x̄, s, and σ. Verify your hand calculation against that output before you move to the median.
Median of Grouped Data
The median of grouped data is the value at which the cumulative relative frequency reaches 0.5. Because the data are continuous, you cannot simply pick the midpoint of the class that contains the median. You must interpolate within that class.
Interpolation formula: Median = L + [ (0.5n, CF) / f_median ] × w
Where L is the lower limit of the median class, CF is the cumulative frequency before the median class, f_median is the frequency of the median class, and w is the class width.
Work the test-score example. n = 50, so the median position is at 25. The cumulative frequencies are:
- 60-69: CF = 5
- 70-79: CF = 5 + 10 = 15
- 80-89: CF = 15 + 20 = 35 (this is the median class because CF crossed 25)
L = 79.5 (the lower limit after the gap), CF = 15, f_median = 20, w = 10. Median = 79.5 + [(25 - 15) / 20] × 10 = 79.5 + (10/20) × 10 = 79.5 + 5 = 84.5.
A student who picks the midpoint of the median class (84.5) gets the same number here by luck because the class is symmetrical. In uneven distributions the interpolation gives a different, more accurate estimate.
Standard Deviation of Grouped Data
The standard deviation of grouped data follows the same estimation logic as the mean. Use the class midpoints as the data values and the frequencies as weights. The sample standard deviation formula is:
Sample SD: s = √[ Σ f (xm, x̄)² / (n, 1) ]
Population SD: σ = √[ Σ f (xm, x̄)² / n ]
For the test-score example, x̄ = 83.5. Compute the squared deviations:
- 60-69: (64.5-83.5)² × 5 = (19)² × 5 = 361 × 5 = 1805
- 70-79: (74.5-83.5)² × 10 = (9)² × 10 = 81 × 10 = 810
- 80-89: (84.5-83.5)² × 20 = (1)² × 20 = 1 × 20 = 20
- 90-99: (94.5-83.5)² × 15 = (11)² × 15 = 121 × 15 = 1815
Σ f (xm, x̄)² = 1805 + 810 + 20 + 1815 = 4450. For a sample, divide by n, 1 = 49: 4450 / 49 ≈ 90.82. The sample standard deviation is √90.82 ≈ 9.53. For a population, divide by n = 50: 4450 / 50 = 89. The population standard deviation is √89 ≈ 9.43.
Sample vs. Population: the Common Mistake
The research confirms that the most frequent error newcomers make is confusing the two denominators. Using n when the data are a sample understates the true variability. Using n, 1 when the data are the entire population overstates it. The TI-84 calculator asks which you want; Excel distinguishes STDEV.S from STDEV.P. The key is to choose the correct one for grouped data.
Why Grouped Results Are Estimates
Every number derived from a frequency table is an approximation. The class midpoint is a stand-in for every value in that class, which assumes values are evenly spread within the interval. Real data are rarely that neat. A class of 60-69 might have all five scores clustered at 61, not spread across the range. The mean calculated from grouped data shifts by the difference between the true scores and the midpoint.
This estimation error shrinks as class width narrows. A table with classes of width 5 produces better estimates than one with width 20. The trade-off is that more classes mean more arithmetic. For most classroom problems, the grouped estimate is close enough to pass a test, but a researcher publishing a study should return to the raw data if possible.
The median estimate from interpolation is less sensitive to the midpoint assumption because it uses the position within the class, not the average value. The standard deviation of grouped data is the most sensitive to the choice of midpoint because it squares deviations.
Common Questions
Can I calculate the mode from a frequency table?
Yes. The modal class is the one with the highest frequency. In the test-score example, the 80-89 class has f = 20, so that is the modal class. A more precise mode requires interpolation within that class, but many textbooks stop at the modal class.
Why does my calculator give a different standard deviation than my hand calculation?
Two common causes. First, you may have used n where the calculator expects n, 1, or the reverse. Second, the calculator may be using a different rounding of the mean. Recompute the mean to at least two decimal places and recheck the squared deviations.
What if the class intervals are not equal width?
Unequal class widths complicate the formulas. The mean formula still works (use each class midpoint and its frequency). The median interpolation must use the actual width of the median class. The standard deviation formula still applies, but the estimate is weaker because the midpoint assumption is less reliable for wide classes.
Is the standard deviation of grouped data biased?
Yes, in two ways. The sample variance is unbiased for the population variance when using n, 1, but the sample standard deviation (the square root) is biased low. That bias is separate from the estimation error introduced by grouping. The grouped estimate is further biased by the midpoint assumption, and the direction depends on the true distribution within each class.
Can I use this method for ordinal data?
No. Ordinal data have order but no consistent numerical distance between categories. Assigning midpoints to ordered categories (like 'agree', 'neutral', 'disagree') creates a false sense of equal spacing. Use the median for ordinal data, not the mean or standard deviation.