10.7Unit 10 Review
A comprehensive review of all Unit 10 topics — scatter plots, correlation, lines of best fit, writing and interpreting linear models, residuals, residual plots, and the correlation coefficient — followed by a full unit assessment.
Why This Matters
Two-variable statistics is one of the most tested data analysis topics on the SAT and ACT. This review ties together every skill from Unit 10 — scatter plots, correlation, lines of best fit, residuals, and the correlation coefficient — so you can approach any statistics question with confidence.
Workbook
Lesson, vocabulary, worked examples, and practice problems.
Unit 10 Review
Two-Variable Statistics — Chapters 01 through 06
In this unit you learned to create and interpret scatter plots, describe and measure correlation, draw and write equations for lines of best fit, make predictions using linear models, calculate and interpret residuals, and use the correlation coefficient to evaluate the strength of a linear relationship. This chapter reviews every key concept and prepares you for the unit assessment.
Use this review to consolidate your understanding of all Unit 10 topics. Work through the concept summaries, then test yourself with the mixed practice problems. Pay special attention to interpreting slope and y-intercept in context, distinguishing interpolation from extrapolation, and reading residual plots.
Concept Summaries
Scatter Plots & Association
- Plot ordered pairs (x, y) on a coordinate plane.
- Describe direction: positive, negative, or no association.
- Describe form: linear or nonlinear.
- Describe strength: strong, moderate, or weak.
- Identify and describe outliers.
Correlation
- Positive correlation: both variables increase together.
- Negative correlation: one increases as the other decreases.
- No correlation: no clear pattern.
- Correlation ≠ causation — a third variable may explain the relationship.
- Lurking variables can create misleading correlations.
Lines of Best Fit
- A line of best fit (trend line) models the linear trend in data.
- Roughly equal numbers of points above and below the line.
- Use the line to make predictions within (interpolation) or beyond (extrapolation) the data range.
- Extrapolation is less reliable — use with caution.
Writing Linear Models
- Write the equation ŷ = mx + b using two points on the line.
- Slope m = rate of change: "for each 1-unit increase in x, y changes by m."
- y-intercept b = predicted value of y when x = 0.
- Always interpret slope and intercept in the context of the problem.
- Substitute x-values to make predictions.
Residuals & Residual Plots
- Residual = actual y − predicted ŷ.
- Positive residual: actual value is above the line.
- Negative residual: actual value is below the line.
- Sum of all residuals ≈ 0 for a good linear model.
- Random scatter in a residual plot → linear model is appropriate.
- A curved pattern in a residual plot → a nonlinear model may be better.
Correlation Coefficient (r)
- r is between −1 and +1.
- Sign of r matches the direction of association.
- |r| close to 1 → strong linear relationship.
- |r| close to 0 → weak or no linear relationship.
- r² = proportion of variation in y explained by x.
- r only measures linear relationships; r = 0 does not mean no relationship.
Visual Reference
Positive Linear Association
Negative Linear Association
No Association
Random Residual Plot (Linear Model OK)
Correlation Coefficient (r) Quick Reference
r = −1
Perfect Neg.
−1 < r ≤ −0.7
Strong Neg.
−0.7 < r ≤ −0.3
Moderate Neg.
−0.3 < r < +0.3
Weak / None
+0.3 ≤ r < +0.7
Moderate Pos.
+0.7 ≤ r ≤ +1
Strong Pos.
Guided Practice Video: Unit 10 Review
Watch this full Unit 10 review covering scatter plots, correlation, lines of best fit, writing equations for lines of best fit, residuals, and the correlation coefficient before tackling the mixed review problems below.
Video by Sang Real Math
Watch on YouTube ↗Mixed Review — Worked Examples
A scatter plot shows hours of TV watched per day (x) and reading test scores (y). The association is negative and strong. What does this mean in context?
Negative association: as hours of TV increase, reading scores tend to decrease.
Strong association: the data points cluster closely around the trend line.
A line of best fit for a data set is ŷ = 2.5x + 10, where x = months of training and y = miles run per week. Interpret the slope and y-intercept.
Slope = 2.5: for each additional month of training, the predicted miles run per week increases by 2.5.
y-intercept = 10: at the start of training (0 months), the predicted miles run per week is 10.
Using ŷ = 2.5x + 10, predict the miles run per week after 8 months. Is this interpolation or extrapolation if the data range is 1–12 months?
Substitute x = 8: ŷ = 2.5(8) + 10 = 20 + 10 = 30.
x = 8 is within the data range of 1–12 months → interpolation.
A data point has an actual value of y = 34 and a predicted value of ŷ = 30. Calculate and interpret the residual.
Residual = actual − predicted = 34 − 30 = +4.
Positive residual: the actual value is 4 units above the line of best fit.
A study reports r = −0.88 between daily sugar intake (grams) and energy levels (1–10 scale). Interpret r and calculate r².
Sign: negative → as sugar intake increases, energy levels tend to decrease.
|r| = 0.88 → strong linear relationship.
r² = (−0.88)² = 0.7744 ≈ 77.4%.
A residual plot for a linear model shows a clear U-shaped curve. What does this tell you about the model?
A curved pattern in a residual plot means the residuals are not randomly scattered.
This indicates the linear model is not appropriate for the data.
Guided Review Problems
A scatter plot shows a positive, moderate, linear association between x and y. Describe what you would expect the data to look like and estimate a reasonable value of r.
Hint: Moderate positive means points trend upward but are somewhat spread out. r would be between +0.3 and +0.7.
The line of best fit for a data set is ŷ = −3x + 95, where x = temperature (°F) and y = hot chocolate sales. Predict sales when the temperature is 20°F. Is this interpolation or extrapolation if the data range is 30–70°F?
Hint: Substitute x = 20 into the equation. Then check whether 20 is inside or outside the range 30–70.
Five residuals for a linear model are: +2, −1, +3, −2, −1. Do these residuals suggest the linear model is a good fit? Explain.
Hint: Check whether the residuals are randomly scattered (positive and negative, no pattern). Also check whether their sum is close to 0.
Two variables have r = 0.95. A researcher concludes that one variable causes the other. Is this conclusion valid? Explain.
Hint: Recall the key principle: correlation does not imply causation. A lurking variable or coincidence could explain the relationship.