9.5Comparing Distributions
Compare multiple data sets using measures of center, spread, box plots, dot plots, and histograms. Draw conclusions and make data-based decisions.
Why This Matters
Comparing distributions is how scientists and analysts draw conclusions from data. Whether you're comparing test scores, climate data, or experimental results, this skill is central to AP Statistics and every data-driven field.
Workbook
Lesson, vocabulary, worked examples, and practice problems.
Essential Question
How can we compare multiple data distributions to make meaningful conclusions?
Lesson Overview
Comparing distributions is one of the most important skills in statistics. When we have two or more data sets, we rarely want to analyze them in isolation — we want to know: Which group performed better? Which is more consistent? Are the distributions similar or different? To answer these questions, we compare distributions across four dimensions: center (where is the typical value?), spread (how variable is the data?), shape (is it symmetric or skewed?), and outliers (are there unusual values?). We can compare distributions using numerical summaries (mean, median, IQR, MAD) or visual displays (box plots, histograms, dot plots). Strong statistical conclusions always reference specific numbers and explain what those numbers mean in the real-world context of the problem.
1. Compare Center
Find the mean and/or median of each data set. State which is higher and by how much. Explain what this means in context.
2. Compare Spread
Find the range, IQR, and/or MAD of each data set. State which has greater spread. Explain what this means (more/less consistent, more/less variable).
3. Compare Shape
Describe the shape of each distribution (symmetric, skewed left, skewed right). Note whether the shapes are similar or different.
4. Check for Outliers
Identify any outliers in each data set. Note how outliers affect the mean and range. Use median and IQR when outliers are present.
5. Overall Conclusion
Synthesize all four dimensions into a clear, evidence-based conclusion. Reference specific numbers. Explain what the comparison means in the real-world context.
Use Mean when…
- The distribution is symmetric
- There are no outliers
- You want to account for every value equally
Use Median when…
- The distribution is skewed
- Outliers are present
- You want a resistant measure of center
Two Box Plots — Side by Side
Group A
Group B
Two Histograms — Same Scale
Class A
Class B
Two Dot Plots — Same Scale
Team A
Team B
Same Center, Different Spread
Set A — Median=50, IQR=20 (more spread)
Set B — Median=50, IQR=6 (less spread)
Both sets have the same median (50), but Set A is much more spread out (IQR=20) while Set B is tightly clustered (IQR=6). Same center ≠ same distribution.
Symmetric vs. Skewed
Symmetric
Median centered; equal whiskers. Mean ≈ Median.
Right-Skewed
Long right tail. Mean > Median.
Left-Skewed
Long left tail. Mean < Median.
Data With and Without Outliers
No Outliers — Range=40, Mean≈30
With Outlier (75) — Range=65, Mean pulled higher
The outlier (75) dramatically increases the range (40→65) and pulls the mean upward, but the median and IQR remain unchanged. This is why median and IQR are preferred for skewed data or data with outliers.
Worked Examples
Compare the two box plots:\nGroup A: Min=20, Q1=35, Median=50, Q3=65, Max=80\nGroup B: Min=30, Q1=55, Median=70, Q3=78, Max=90
Step 1 — Compare Center: Group A median=50; Group B median=70. Group B has a higher typical value by 20 points.
Step 2 — Compare Spread: Group A IQR=65−35=30; Group B IQR=78−55=23. Group A has more variability in the middle 50% (IQR 30 vs. 23). Group A range=60; Group B range=60. Equal total range.
Step 3 — Compare Shape: Group A — median (50) is centered in the box → roughly symmetric. Group B — median (70) is closer to Q3 (78) → left-skewed.
Step 4 — Outliers: No outliers visible in either box plot.
Conclusion: Group B has a higher typical value (median 70 vs. 50). Group A has more variability in the middle 50% (IQR 30 vs. 23). Group B is left-skewed; Group A is roughly symmetric.
Diagram — Example 1
Group A — Median=50, IQR=30 (symmetric)
Group B — Median=70, IQR=23 (left-skewed)
Group A
Median = 50 · IQR = 30 · Range = 60
Median centered in box → symmetric
Group B
Median = 70 · IQR = 23 · Range = 60
Median near Q3 → left-skewed
Compare the two histograms:\nClass A: 50–59: 2, 60–69: 5, 70–79: 8, 80–89: 6, 90–99: 3 (n=24)\nClass B: 50–59: 1, 60–69: 2, 70–79: 4, 80–89: 9, 90–99: 8 (n=24)
Step 1 — Shape: Class A peaks in 70–79 and is roughly symmetric. Class B peaks in 80–89 and is left-skewed (most scores are high, with a tail toward lower scores).
Step 2 — Center: Class A mode interval = 70–79 (midpoint ≈ 74.5). Class B mode interval = 80–89 (midpoint ≈ 84.5). Class B has a higher center.
Step 3 — Spread: Class A scores range from 50–99 (range ≈ 49). Class B also ranges from 50–99 (range ≈ 49). Similar total range, but Class B is concentrated in the upper ranges.
Step 4 — Conclusion: Class B performed better overall — more students scored in the 80–99 range (17 out of 24 = 71%) compared to Class A (9 out of 24 = 38%). Class A has a more even distribution across all score ranges.
Diagram — Example 2: Side-by-Side Histograms (Same Scale)
Class A — Roughly Symmetric
Class B — Left-Skewed
| Interval | Class A (n) | Class A (%) | Class B (n) | Class B (%) |
|---|---|---|---|---|
| 50–59 | 2 | 8% | 1 | 4% |
| 60–69 | 5 | 21% | 2 | 8% |
| 70–79 | 8 | 33% | 4 | 17% |
| 80–89 | 6 | 25% | 9 | 38% |
| 90–99 | 3 | 13% | 8 | 33% |
| 80–99 total | 9 | 38% | 17 | 71% |
Class A — Symmetric
Peak interval: 70–79 (8 students, 33%)
Scores spread evenly across all ranges
80–99: 9/24 = 38%
Shape: bell-curve, mean ≈ median ≈ 74
Class B — Left-Skewed
Peak interval: 80–89 (9 students, 38%)
Scores concentrated in upper ranges
80–99: 17/24 = 71%
Shape: tail toward lower scores, mean < median
Compare the two dot plots:\nTeam A: 3, 4, 4, 5, 5, 5, 6, 6, 7, 8\nTeam B: 1, 2, 5, 5, 5, 5, 5, 8, 9, 10
Step 1 — Order and find medians: Team A ordered: 3,4,4,5,5,5,6,6,7,8. Median=(5+5)/2=5. Team B ordered: 1,2,5,5,5,5,5,8,9,10. Median=(5+5)/2=5.
Step 2 — Compare Center: Both teams have median=5. Same typical value.
Step 3 — Compare Spread: Team A range=8−3=5. Team B range=10−1=9. Team B has much greater total spread.
Step 4 — Shape and Clusters: Team A clusters tightly around 4–6 (consistent). Team B has a cluster at 5 but values spread from 1 to 10 (inconsistent).
Step 5 — Conclusion: Both teams have the same median (5), but Team B is far more variable (range=9 vs. 5). Team A is more consistent and predictable. Team B has extreme values on both ends.
Diagram — Example 3
Team A — Consistent (Range=5)
Team B — Variable (Range=9)
Team A
Median = 5 · Range = 5
Tight cluster around 4–6 → consistent
Team B
Median = 5 · Range = 9
Spread 1–10 with cluster at 5 → variable
Two job candidates' interview scores (out of 10) over 8 rounds:\nCandidate A: 6, 7, 7, 8, 8, 8, 9, 9\nCandidate B: 4, 5, 7, 8, 8, 9, 10, 10\nWhich candidate should be hired if consistency matters most?
Order both: A: 6,7,7,8,8,8,9,9. B: 4,5,7,8,8,9,10,10.
Candidate A: Mean=(6+7+7+8+8+8+9+9)/8=62/8=7.75. Median=(8+8)/2=8. Range=9−6=3. Q1=7, Q3=9, IQR=2.
Candidate B: Mean=(4+5+7+8+8+9+10+10)/8=61/8=7.625. Median=(8+8)/2=8. Range=10−4=6. Q1=6, Q3=9.5, IQR=3.5.
Compare: Nearly identical means (7.75 vs. 7.625) and medians (both 8). But Candidate A has a much smaller range (3 vs. 6) and IQR (2 vs. 3.5).
Conclusion: If consistency matters most, hire Candidate A. Their scores are tightly clustered (range=3, IQR=2) with no low scores. Candidate B has higher highs (10) but also lower lows (4), making them less predictable.
Diagram — Example 4
Candidate A — Median=8, IQR=2, Range=3 (consistent)
Candidate B — Median=8, IQR=3.5, Range=6 (variable)
Candidate A ✓ More Consistent
Median=8 · IQR=2 · Range=3
Narrow box → tight cluster → reliable
Candidate B — More Variable
Median=8 · IQR=3.5 · Range=6
Wider box + longer whiskers → unpredictable
Interpret center and spread:\nData Set X: Mean=45, Median=44, IQR=12, Range=30\nData Set Y: Mean=45, Median=38, IQR=10, Range=55\nWhat do these statistics tell you about each distribution?
Set X: Mean (45) ≈ Median (44) → roughly symmetric distribution. IQR=12, Range=30 → moderate spread.
Set Y: Mean (45) >> Median (38) → right-skewed distribution (mean pulled up by high values). IQR=10 (smaller than X), but Range=55 (much larger than X).
Interpretation: Set Y has a large outlier or extreme high value that pulls the mean far above the median and inflates the range. The middle 50% of Set Y (IQR=10) is actually less spread than Set X (IQR=12), but the tail is much longer.
Best measure of center: Set X → mean (symmetric). Set Y → median (skewed, outlier present).
Conclusion: Both sets have the same mean, but they are very different distributions. Set X is symmetric and moderately spread. Set Y is right-skewed with an extreme upper tail.
Diagram — Example 5
Set X — Mean=45, Median=44 (symmetric, IQR=12)
Set Y — Mean=45, Median=38 (right-skewed, IQR=10, Range=55)
Set X — Symmetric
Mean=45 ≈ Median=44 · IQR=12 · Range=30
Use mean as measure of center
Set Y — Right-Skewed
Mean=45 > Median=38 · IQR=10 · Range=55
Long right tail pulls mean up → use median
Notice: Set Y has a smaller IQR (10 vs. 12) but a much larger range (55 vs. 30). The middle 50% is tighter, but an extreme high value creates a long right tail — exactly why range alone is misleading.
Guided Practice
Guided Practice Video: Comparing Distributions
Review how to compare two or more distributions using shape, center, and spread — including side-by-side box plots and back-to-back dot plots — before completing the guided problems below.
Video by Sang Real Math
Watch on YouTube ↗Two box plots: Set A has Min=10, Q1=25, Median=40, Q3=55, Max=70. Set B has Min=20, Q1=45, Median=60, Q3=70, Max=80. Compare the center and spread of both sets.
Hint: Find IQR for each (Q3−Q1). Compare medians for center. Compare IQRs for spread. Which set has higher typical values? Which is more variable?
Data Set P: 5, 8, 10, 12, 15, 18, 20, 22, 25, 28. Data Set Q: 5, 5, 5, 10, 15, 20, 25, 30, 30, 30. Compare the center, spread, and shape of both sets.
Hint: Find the median and IQR for each. Describe the shape of each distribution. Are they symmetric or skewed? Do they have the same center?
Two histograms show test scores. Histogram A peaks in the 70–79 range and is symmetric. Histogram B peaks in the 90–99 range and is left-skewed. Which class performed better? Which has more variability?
Hint: The peak of a histogram is near the mode/center. Left-skewed means most scores are high with a tail toward lower scores. Compare the centers and shapes.
Dot Plot X: 2, 3, 3, 4, 4, 4, 5, 5, 6, 7. Dot Plot Y: 1, 2, 4, 4, 4, 4, 4, 6, 8, 9. Both have the same median. Which has greater spread? Which is more consistent?
Hint: Find the median and range for each. Compare the IQR. Which dot plot has values more tightly clustered around the center?
Group A: Mean=65, Median=64, IQR=15. Group B: Mean=65, Median=55, IQR=12. A student says both groups performed equally because they have the same mean. What is wrong with this conclusion?
Hint: Look at the medians — they are very different. What does it mean when the mean is much higher than the median? Which group has a skewed distribution?
Key Vocabulary
Distribution
The pattern of how data values are spread across possible values. Described by shape, center, and spread.
Center
A single value that represents the typical or middle value of a data set. Measured by mean or median.
Spread
How far data values are from each other or from the center. Measured by range, IQR, or MAD.
Variability
The degree to which data values differ from each other. High variability = data is spread out. Low variability = data is clustered together.
Shape
The overall pattern of a distribution: symmetric, skewed left, skewed right, uniform, or bimodal.
Skewness
The degree of asymmetry in a distribution. Right-skewed: long right tail. Left-skewed: long left tail. Symmetric: balanced on both sides.
Symmetric Distribution
A distribution where the left and right sides are mirror images. Mean ≈ median. Box plot: median centered in box, equal whiskers.
Outlier
A data value that is unusually far from the rest of the data. Identified using the IQR fence method: below Q1 − 1.5(IQR) or above Q3 + 1.5(IQR).
Cluster
A group of data values that are close together in a dot plot or histogram, indicating a concentration of data in that region.
Gap
A region in a dot plot or histogram where no data values appear, indicating a break in the distribution.
Mean
The arithmetic average of a data set. Sensitive to outliers. Best for symmetric distributions.
Median
The middle value of an ordered data set. Resistant to outliers. Best for skewed distributions.
IQR
Interquartile Range = Q3 − Q1. Measures the spread of the middle 50% of data. Resistant to outliers.
MAD
Mean Absolute Deviation. The average distance of each data value from the mean. Measures typical spread around the mean.
Interactive Practice — 5 Questions
Two data sets have the same mean of 50. Which statement is always true?
Box Plot A has Median=50, IQR=30. Box Plot B has Median=50, IQR=8. Which data set is more consistent?
A distribution has Mean=80 and Median=65. The distribution is most likely:
When comparing two skewed data sets, which pair of statistics is MOST appropriate?
Two box plots are drawn on the same scale. Box Plot X has a wider box than Box Plot Y. This means:
Independent Practice
Independent Practice
Box plots: Set A has Min=15, Q1=30, Median=45, Q3=60, Max=75. Set B has Min=25, Q1=50, Median=65, Q3=75, Max=85. Find the IQR for each. Compare the center and spread in 2–3 sentences.
Data Set M: 10, 15, 20, 25, 30, 35, 40, 45, 50, 55. Data Set N: 10, 10, 10, 25, 30, 35, 50, 55, 55, 55. Find the median and IQR for each. Compare the distributions — which is more consistent and why?
Two dot plots: Plot A has values 4, 5, 5, 6, 6, 6, 7, 7, 8. Plot B has values 1, 3, 5, 6, 6, 7, 9, 11, 13. Compare the center, spread, and shape of both plots.
Group X: Mean=50, Median=49, IQR=12, Range=30. Group Y: Mean=50, Median=42, IQR=10, Range=55. Explain why these two groups are NOT equally distributed even though they have the same mean.
Two box plots both have Median=60. Box A has IQR=25; Box B has IQR=8. Which data set is more consistent? Write a complete 2–3 sentence comparison including center, spread, and what you would conclude.
Common Mistakes
Comparing only the means without considering spread — two distributions can have the same mean but very different variability.
Always compare both center (mean or median) AND spread (range, IQR, or MAD) when describing distributions.
Describing a distribution as 'better' or 'worse' without context — e.g., 'Group A is better because its mean is higher'.
Whether a higher or lower value is better depends on the context. Always state what the numbers represent.
Forgetting to state context when comparing — just saying 'Group A has a higher median' without explaining what that means.
Always connect statistics to the real-world situation: 'Group A's median test score is 8 points higher than Group B's.'
Confusing skewness direction — saying data is right-skewed when the tail points left.
Right-skewed (positive skew): tail points right, mean > median. Left-skewed (negative skew): tail points left, mean < median.
Math Tips
Always compare in context — state what the numbers mean, not just the numbers themselves.
Use specific numbers: "Group A IQR (15) is larger than Group B IQR (8), so Group A has more variability."
Compare all four dimensions: center, spread, shape, and outliers. A mean-only comparison is incomplete.
Larger IQR = more spread in the middle 50%. Larger range = more total spread. They measure different things.
Same center ≠ same distribution. Two data sets can have identical means but very different spreads and shapes.
Use median for skewed data or data with outliers. Use mean for symmetric data with no outliers.