Unit 9 · Lesson 9.5

9.5Comparing Distributions

Compare multiple data sets using measures of center, spread, box plots, dot plots, and histograms. Draw conclusions and make data-based decisions.

Why This Matters

Comparing distributions is how scientists and analysts draw conclusions from data. Whether you're comparing test scores, climate data, or experimental results, this skill is central to AP Statistics and every data-driven field.

Workbook

Lesson, vocabulary, worked examples, and practice problems.

Essential Question

How can we compare multiple data distributions to make meaningful conclusions?

Lesson Overview

Comparing distributions is one of the most important skills in statistics. When we have two or more data sets, we rarely want to analyze them in isolation — we want to know: Which group performed better? Which is more consistent? Are the distributions similar or different? To answer these questions, we compare distributions across four dimensions: center (where is the typical value?), spread (how variable is the data?), shape (is it symmetric or skewed?), and outliers (are there unusual values?). We can compare distributions using numerical summaries (mean, median, IQR, MAD) or visual displays (box plots, histograms, dot plots). Strong statistical conclusions always reference specific numbers and explain what those numbers mean in the real-world context of the problem.

1. Compare Center

Find the mean and/or median of each data set. State which is higher and by how much. Explain what this means in context.

2. Compare Spread

Find the range, IQR, and/or MAD of each data set. State which has greater spread. Explain what this means (more/less consistent, more/less variable).

3. Compare Shape

Describe the shape of each distribution (symmetric, skewed left, skewed right). Note whether the shapes are similar or different.

4. Check for Outliers

Identify any outliers in each data set. Note how outliers affect the mean and range. Use median and IQR when outliers are present.

5. Overall Conclusion

Synthesize all four dimensions into a clear, evidence-based conclusion. Reference specific numbers. Explain what the comparison means in the real-world context.

Use Mean when…

  • The distribution is symmetric
  • There are no outliers
  • You want to account for every value equally

Use Median when…

  • The distribution is skewed
  • Outliers are present
  • You want a resistant measure of center

Two Box Plots — Side by Side

Group A

40
50
60
70
80
90

Group B

40
50
60
70
80
90
Reading: Group B has a higher median (80 vs. 65) and higher Q1/Q3 — better overall performance. Group A has a wider box (IQR=20 vs. 18) and longer left whisker — more variability and lower scores on the left tail.

Two Histograms — Same Scale

Class A

2
5
8
6
3
50–59
60–69
70–79
80–89
90–99

Class B

1
2
4
9
8
50–59
60–69
70–79
80–89
90–99
Reading: Class A peaks in the 70–79 range (roughly symmetric). Class B peaks in the 80–89 range and is left-skewed — most students scored higher. Class B has a higher center; Class A has more spread across the lower ranges.

Two Dot Plots — Same Scale

Team A

1
2
3
4
5
6
7
8
9
10

Team B

1
2
3
4
5
6
7
8
9
10
Reading: Team A clusters tightly around 5–6 (low variability, consistent). Team B is spread across 1–10 with a cluster at 5 (high variability, inconsistent). Both have similar medians (~5) but very different spreads.

Same Center, Different Spread

Set A — Median=50, IQR=20 (more spread)

20
30
40
50
60
70
80

Set B — Median=50, IQR=6 (less spread)

20
30
40
50
60
70
80

Both sets have the same median (50), but Set A is much more spread out (IQR=20) while Set B is tightly clustered (IQR=6). Same center ≠ same distribution.

Symmetric vs. Skewed

Symmetric

10
25
40
55
70

Median centered; equal whiskers. Mean ≈ Median.

Right-Skewed

10
20
40
60
80

Long right tail. Mean > Median.

Left-Skewed

10
30
50
70
80

Long left tail. Mean < Median.

Data With and Without Outliers

No Outliers — Range=40, Mean≈30

10
20
30
40
50
60
70

With Outlier (75) — Range=65, Mean pulled higher

10
20
30
40
50
60
70

The outlier (75) dramatically increases the range (40→65) and pulls the mean upward, but the median and IQR remain unchanged. This is why median and IQR are preferred for skewed data or data with outliers.

Worked Examples

Example 1

Compare the two box plots:\nGroup A: Min=20, Q1=35, Median=50, Q3=65, Max=80\nGroup B: Min=30, Q1=55, Median=70, Q3=78, Max=90

Step 1 — Compare Center: Group A median=50; Group B median=70. Group B has a higher typical value by 20 points.

Step 2 — Compare Spread: Group A IQR=65−35=30; Group B IQR=78−55=23. Group A has more variability in the middle 50% (IQR 30 vs. 23). Group A range=60; Group B range=60. Equal total range.

Step 3 — Compare Shape: Group A — median (50) is centered in the box → roughly symmetric. Group B — median (70) is closer to Q3 (78) → left-skewed.

Step 4 — Outliers: No outliers visible in either box plot.

Conclusion: Group B has a higher typical value (median 70 vs. 50). Group A has more variability in the middle 50% (IQR 30 vs. 23). Group B is left-skewed; Group A is roughly symmetric.

Answer:Group B has a higher center (median=70 vs. 50). Group A has greater spread (IQR=30 vs. 23). Group B is left-skewed; Group A is symmetric.

Diagram — Example 1

Group A — Median=50, IQR=30 (symmetric)

20
30
40
50
60
70
80
90

Group B — Median=70, IQR=23 (left-skewed)

20
30
40
50
60
70
80
90

Group A

Median = 50 · IQR = 30 · Range = 60

Median centered in box → symmetric

Group B

Median = 70 · IQR = 23 · Range = 60

Median near Q3 → left-skewed

Example 2

Compare the two histograms:\nClass A: 50–59: 2, 60–69: 5, 70–79: 8, 80–89: 6, 90–99: 3 (n=24)\nClass B: 50–59: 1, 60–69: 2, 70–79: 4, 80–89: 9, 90–99: 8 (n=24)

Step 1 — Shape: Class A peaks in 70–79 and is roughly symmetric. Class B peaks in 80–89 and is left-skewed (most scores are high, with a tail toward lower scores).

Step 2 — Center: Class A mode interval = 70–79 (midpoint ≈ 74.5). Class B mode interval = 80–89 (midpoint ≈ 84.5). Class B has a higher center.

Step 3 — Spread: Class A scores range from 50–99 (range ≈ 49). Class B also ranges from 50–99 (range ≈ 49). Similar total range, but Class B is concentrated in the upper ranges.

Step 4 — Conclusion: Class B performed better overall — more students scored in the 80–99 range (17 out of 24 = 71%) compared to Class A (9 out of 24 = 38%). Class A has a more even distribution across all score ranges.

Answer:Class B has a higher center and more students in the 80–99 range (71% vs. 38%). Class A is roughly symmetric; Class B is left-skewed.

Diagram — Example 2: Side-by-Side Histograms (Same Scale)

Class A — Roughly Symmetric

02468102586350–5960–6970–7980–8990–9980–99 ↑

Class B — Left-Skewed

02468101249850–5960–6970–7980–8990–9980–99 ↑
IntervalClass A (n)Class A (%)Class B (n)Class B (%)
50–5928%14%
60–69521%28%
70–79833%417%
80–89625%938%
90–99313%833%
80–99 total938%1771%

Class A — Symmetric

Peak interval: 70–79 (8 students, 33%)

Scores spread evenly across all ranges

80–99: 9/24 = 38%

Shape: bell-curve, mean ≈ median ≈ 74

Class B — Left-Skewed

Peak interval: 80–89 (9 students, 38%)

Scores concentrated in upper ranges

80–99: 17/24 = 71%

Shape: tail toward lower scores, mean < median

Key insight: Class B has nearly twice as many students scoring 80–99 (71% vs. 38%). Despite both classes having the same score range (50–99), Class B's distribution is shifted right — its center is higher and its tail extends left toward the lower scores (left-skewed).
Example 3

Compare the two dot plots:\nTeam A: 3, 4, 4, 5, 5, 5, 6, 6, 7, 8\nTeam B: 1, 2, 5, 5, 5, 5, 5, 8, 9, 10

Step 1 — Order and find medians: Team A ordered: 3,4,4,5,5,5,6,6,7,8. Median=(5+5)/2=5. Team B ordered: 1,2,5,5,5,5,5,8,9,10. Median=(5+5)/2=5.

Step 2 — Compare Center: Both teams have median=5. Same typical value.

Step 3 — Compare Spread: Team A range=8−3=5. Team B range=10−1=9. Team B has much greater total spread.

Step 4 — Shape and Clusters: Team A clusters tightly around 4–6 (consistent). Team B has a cluster at 5 but values spread from 1 to 10 (inconsistent).

Step 5 — Conclusion: Both teams have the same median (5), but Team B is far more variable (range=9 vs. 5). Team A is more consistent and predictable. Team B has extreme values on both ends.

Answer:Same median (5), but Team B has much greater spread (range=9 vs. 5). Team A is more consistent; Team B is highly variable.

Diagram — Example 3

Team A — Consistent (Range=5)

1
2
3
4
5
6
7
8
9
10

Team B — Variable (Range=9)

1
2
3
4
5
6
7
8
9
10

Team A

Median = 5 · Range = 5

Tight cluster around 4–6 → consistent

Team B

Median = 5 · Range = 9

Spread 1–10 with cluster at 5 → variable

Example 4

Two job candidates' interview scores (out of 10) over 8 rounds:\nCandidate A: 6, 7, 7, 8, 8, 8, 9, 9\nCandidate B: 4, 5, 7, 8, 8, 9, 10, 10\nWhich candidate should be hired if consistency matters most?

Order both: A: 6,7,7,8,8,8,9,9. B: 4,5,7,8,8,9,10,10.

Candidate A: Mean=(6+7+7+8+8+8+9+9)/8=62/8=7.75. Median=(8+8)/2=8. Range=9−6=3. Q1=7, Q3=9, IQR=2.

Candidate B: Mean=(4+5+7+8+8+9+10+10)/8=61/8=7.625. Median=(8+8)/2=8. Range=10−4=6. Q1=6, Q3=9.5, IQR=3.5.

Compare: Nearly identical means (7.75 vs. 7.625) and medians (both 8). But Candidate A has a much smaller range (3 vs. 6) and IQR (2 vs. 3.5).

Conclusion: If consistency matters most, hire Candidate A. Their scores are tightly clustered (range=3, IQR=2) with no low scores. Candidate B has higher highs (10) but also lower lows (4), making them less predictable.

Answer:Hire Candidate A for consistency. Same median (8), but A has smaller range (3 vs. 6) and IQR (2 vs. 3.5). A is more reliable.

Diagram — Example 4

Candidate A — Median=8, IQR=2, Range=3 (consistent)

4
5
6
7
8
9
10

Candidate B — Median=8, IQR=3.5, Range=6 (variable)

4
5
6
7
8
9
10

Candidate A ✓ More Consistent

Median=8 · IQR=2 · Range=3

Narrow box → tight cluster → reliable

Candidate B — More Variable

Median=8 · IQR=3.5 · Range=6

Wider box + longer whiskers → unpredictable

Example 5

Interpret center and spread:\nData Set X: Mean=45, Median=44, IQR=12, Range=30\nData Set Y: Mean=45, Median=38, IQR=10, Range=55\nWhat do these statistics tell you about each distribution?

Set X: Mean (45) ≈ Median (44) → roughly symmetric distribution. IQR=12, Range=30 → moderate spread.

Set Y: Mean (45) >> Median (38) → right-skewed distribution (mean pulled up by high values). IQR=10 (smaller than X), but Range=55 (much larger than X).

Interpretation: Set Y has a large outlier or extreme high value that pulls the mean far above the median and inflates the range. The middle 50% of Set Y (IQR=10) is actually less spread than Set X (IQR=12), but the tail is much longer.

Best measure of center: Set X → mean (symmetric). Set Y → median (skewed, outlier present).

Conclusion: Both sets have the same mean, but they are very different distributions. Set X is symmetric and moderately spread. Set Y is right-skewed with an extreme upper tail.

Answer:Set X is symmetric (mean≈median). Set Y is right-skewed (mean>>median) with an extreme upper tail. Use median for Set Y. Same mean ≠ same distribution.

Diagram — Example 5

Set X — Mean=45, Median=44 (symmetric, IQR=12)

20
30
40
50
60
70
80
90

Set Y — Mean=45, Median=38 (right-skewed, IQR=10, Range=55)

20
30
40
50
60
70
80
90

Set X — Symmetric

Mean=45 ≈ Median=44 · IQR=12 · Range=30

Use mean as measure of center

Set Y — Right-Skewed

Mean=45 > Median=38 · IQR=10 · Range=55

Long right tail pulls mean up → use median

Notice: Set Y has a smaller IQR (10 vs. 12) but a much larger range (55 vs. 30). The middle 50% is tighter, but an extreme high value creates a long right tail — exactly why range alone is misleading.

Guided Practice

Guided Practice Video: Comparing Distributions

Review how to compare two or more distributions using shape, center, and spread — including side-by-side box plots and back-to-back dot plots — before completing the guided problems below.

Video by Sang Real Math

Watch on YouTube ↗
Guided Problem 1

Two box plots: Set A has Min=10, Q1=25, Median=40, Q3=55, Max=70. Set B has Min=20, Q1=45, Median=60, Q3=70, Max=80. Compare the center and spread of both sets.

Hint: Find IQR for each (Q3−Q1). Compare medians for center. Compare IQRs for spread. Which set has higher typical values? Which is more variable?

Guided Problem 2

Data Set P: 5, 8, 10, 12, 15, 18, 20, 22, 25, 28. Data Set Q: 5, 5, 5, 10, 15, 20, 25, 30, 30, 30. Compare the center, spread, and shape of both sets.

Hint: Find the median and IQR for each. Describe the shape of each distribution. Are they symmetric or skewed? Do they have the same center?

Guided Problem 3

Two histograms show test scores. Histogram A peaks in the 70–79 range and is symmetric. Histogram B peaks in the 90–99 range and is left-skewed. Which class performed better? Which has more variability?

Hint: The peak of a histogram is near the mode/center. Left-skewed means most scores are high with a tail toward lower scores. Compare the centers and shapes.

Guided Problem 4

Dot Plot X: 2, 3, 3, 4, 4, 4, 5, 5, 6, 7. Dot Plot Y: 1, 2, 4, 4, 4, 4, 4, 6, 8, 9. Both have the same median. Which has greater spread? Which is more consistent?

Hint: Find the median and range for each. Compare the IQR. Which dot plot has values more tightly clustered around the center?

Guided Problem 5

Group A: Mean=65, Median=64, IQR=15. Group B: Mean=65, Median=55, IQR=12. A student says both groups performed equally because they have the same mean. What is wrong with this conclusion?

Hint: Look at the medians — they are very different. What does it mean when the mean is much higher than the median? Which group has a skewed distribution?

Key Vocabulary

Distribution

The pattern of how data values are spread across possible values. Described by shape, center, and spread.

Center

A single value that represents the typical or middle value of a data set. Measured by mean or median.

Spread

How far data values are from each other or from the center. Measured by range, IQR, or MAD.

Variability

The degree to which data values differ from each other. High variability = data is spread out. Low variability = data is clustered together.

Shape

The overall pattern of a distribution: symmetric, skewed left, skewed right, uniform, or bimodal.

Skewness

The degree of asymmetry in a distribution. Right-skewed: long right tail. Left-skewed: long left tail. Symmetric: balanced on both sides.

Symmetric Distribution

A distribution where the left and right sides are mirror images. Mean ≈ median. Box plot: median centered in box, equal whiskers.

Outlier

A data value that is unusually far from the rest of the data. Identified using the IQR fence method: below Q1 − 1.5(IQR) or above Q3 + 1.5(IQR).

Cluster

A group of data values that are close together in a dot plot or histogram, indicating a concentration of data in that region.

Gap

A region in a dot plot or histogram where no data values appear, indicating a break in the distribution.

Mean

The arithmetic average of a data set. Sensitive to outliers. Best for symmetric distributions.

Median

The middle value of an ordered data set. Resistant to outliers. Best for skewed distributions.

IQR

Interquartile Range = Q3 − Q1. Measures the spread of the middle 50% of data. Resistant to outliers.

MAD

Mean Absolute Deviation. The average distance of each data value from the mean. Measures typical spread around the mean.

Interactive Practice — 5 Questions

1

Two data sets have the same mean of 50. Which statement is always true?

2

Box Plot A has Median=50, IQR=30. Box Plot B has Median=50, IQR=8. Which data set is more consistent?

3

A distribution has Mean=80 and Median=65. The distribution is most likely:

4

When comparing two skewed data sets, which pair of statistics is MOST appropriate?

5

Two box plots are drawn on the same scale. Box Plot X has a wider box than Box Plot Y. This means:

Independent Practice

Independent Practice

1

Box plots: Set A has Min=15, Q1=30, Median=45, Q3=60, Max=75. Set B has Min=25, Q1=50, Median=65, Q3=75, Max=85. Find the IQR for each. Compare the center and spread in 2–3 sentences.

2

Data Set M: 10, 15, 20, 25, 30, 35, 40, 45, 50, 55. Data Set N: 10, 10, 10, 25, 30, 35, 50, 55, 55, 55. Find the median and IQR for each. Compare the distributions — which is more consistent and why?

3

Two dot plots: Plot A has values 4, 5, 5, 6, 6, 6, 7, 7, 8. Plot B has values 1, 3, 5, 6, 6, 7, 9, 11, 13. Compare the center, spread, and shape of both plots.

4

Group X: Mean=50, Median=49, IQR=12, Range=30. Group Y: Mean=50, Median=42, IQR=10, Range=55. Explain why these two groups are NOT equally distributed even though they have the same mean.

5

Two box plots both have Median=60. Box A has IQR=25; Box B has IQR=8. Which data set is more consistent? Write a complete 2–3 sentence comparison including center, spread, and what you would conclude.

⚠️

Common Mistakes

Comparing only the means without considering spread — two distributions can have the same mean but very different variability.

Always compare both center (mean or median) AND spread (range, IQR, or MAD) when describing distributions.

Describing a distribution as 'better' or 'worse' without context — e.g., 'Group A is better because its mean is higher'.

Whether a higher or lower value is better depends on the context. Always state what the numbers represent.

Forgetting to state context when comparing — just saying 'Group A has a higher median' without explaining what that means.

Always connect statistics to the real-world situation: 'Group A's median test score is 8 points higher than Group B's.'

Confusing skewness direction — saying data is right-skewed when the tail points left.

Right-skewed (positive skew): tail points right, mean > median. Left-skewed (negative skew): tail points left, mean < median.

💡

Math Tips

📌

Always compare in context — state what the numbers mean, not just the numbers themselves.

📌

Use specific numbers: "Group A IQR (15) is larger than Group B IQR (8), so Group A has more variability."

📌

Compare all four dimensions: center, spread, shape, and outliers. A mean-only comparison is incomplete.

📌

Larger IQR = more spread in the middle 50%. Larger range = more total spread. They measure different things.

📌

Same center ≠ same distribution. Two data sets can have identical means but very different spreads and shapes.

📌

Use median for skewed data or data with outliers. Use mean for symmetric data with no outliers.