Linear regression:
slope, intercept and r² from five data points
The fitted line is nothing more than two averages and two sums. Do it once by hand and every software output makes sense.
Calcylator Editorial Team
Updated · 4 min read
What the best-fit line does
Simple linear regression draws the straight line that sits as close as possible to a cloud of points, where closeness is measured by the vertical gap between each point and the line. Square those gaps, add them up, and choose the line that makes that sum the smallest. That is the method of least squares.
The result is an equation y = a + bx. The slope b says how much y changes, on average, for each one-unit increase in x. The intercept a is the predicted y when x is zero, which may or may not have a practical meaning.
The aim is usually either description (how strongly are two things linked) or prediction (what should y be for a new x). The same line serves both, but the cautions differ.
The formulas
- x̄, ȳ:
- the means of the x and y values
- Σ:
- sum over all data points
- a:
- the value of y where the line crosses x = 0
- residual:
- actual y − predicted y
- total sum of squares:
- Σ(y − ȳ)²
The numerator of the slope is the shared variation of x and y, and the denominator is the variation of x alone. Their ratio is how much y moves for each unit of x.
Worked example: study hours and marks
Five students record weekly study hours (x) and their test marks out of 100 (y): (2, 52), (4, 60), (5, 68), (7, 75), (9, 88). Start with the means: x̄ = 27 ÷ 5 = 5.4 and ȳ = 343 ÷ 5 = 68.6.
Σ(x − x̄)(y − ȳ)
148.8
Σ(x − x̄)²
29.2
Slope b
148.8 ÷ 29.2 = 5.096
Intercept a
68.6 − 5.096 × 5.4 = 41.08
Σ(y − ȳ)²
767.2
Fitted line
y = 41.08 + 5.10x
r = 0.994 and r² = 0.988, so about 98.8 percent of the variation in marks is accounted for by hours.
Reading the slope: each extra study hour is associated with about 5.1 additional marks across these five students. A student studying 6 hours is predicted to score 41.08 + 5.096 × 6 = 71.7, and one studying 8 hours to score 81.8.
To confirm the fit, the predicted values are 51.3, 61.5, 66.6, 76.8 and 86.9. The residuals (actual minus predicted) are +0.7, −1.5, +1.4, −1.8 and +1.1. They sum to zero, as least-squares residuals always do.
How to judge whether the line is any good
- Plot the points first. A straight line only suits data that look roughly straight.
- Look at residuals. They should scatter above and below zero without a pattern; a curve means a line is the wrong shape.
- Check r². A high value shows a close fit, but with few points a high r² can occur by chance.
- Check the sample size. Five points are enough to learn the method, not to base decisions on.
Correlation does not establish cause. Students who study more may also differ in motivation, prior knowledge or coaching, and the regression cannot separate those factors. It describes the pattern in the data you gave it.
Predicting outside the data range
A fitted line is trustworthy between the smallest and largest x in the data, and it becomes increasingly fragile outside. Our line says that 15 study hours yields 41.08 + 5.096 × 15 = 117.5 marks, which is impossible on a test marked out of 100.
The straight line was only ever an approximation over the observed range. Real relationships flatten, curve or hit limits. Interpolating between known points is generally safe; extrapolating is a guess.
Residual spread and why the roles of x and y matter
The line is only part of the story. The typical size of the misses is summarised by the standard error of the estimate, found by squaring the residuals, adding them, dividing by n − 2 and taking the square root. In the study example the squared residuals add up to 8.93, so the standard error is √(8.93 ÷ 3) = 1.73 marks. Predictions from the line are therefore usually within a couple of marks for students inside the observed range.
Regression is not symmetric. If you swap the variables and regress hours on marks, the slope becomes Sxy ÷ Syy = 148.8 ÷ 767.2 = 0.194 hours per mark, which is not the reciprocal of 5.10. The line that predicts marks from hours differs from the line that predicts hours from marks, because each minimises vertical gaps in its own variable.
In a spreadsheet the same quantities are available directly: SLOPE and INTERCEPT take the y range followed by the x range, RSQ returns r², and FORECAST or FORECAST.LINEAR predicts for a new x. Checking a hand calculation on five points against these functions is a good way to build trust in the tool before it is used on larger data.
Assumptions and what to do when they fail
Ordinary least squares works best when the relationship is roughly linear, the residuals have similar spread across the range, observations are independent and there are no extreme outliers. A single unusual point can pull the line strongly toward itself.
If a plot shows a curve, try transforming a variable, for instance using the logarithm, or fit a polynomial. If you have several predictors, multiple regression extends the same idea. A calculator or spreadsheet function such as SLOPE and INTERCEPT will return the same numbers as the manual steps above, and it is worth checking one small example by hand to be sure you read the output correctly.
Common questions
How do you calculate the linear regression line?
Compute the means of x and y, then the slope b = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)², and the intercept a = ȳ − b·x̄. The line is y = a + bx. For the five-point study example, b is 5.10 and a is 41.08.
What does the slope of a regression line mean?
The slope is the average change in y for each one-unit increase in x. A slope of 5.1 in the study-hours example means each additional hour is associated with about 5.1 more marks. It describes association, not necessarily cause.
What is r-squared in linear regression?
R-squared is the share of the variation in y that the line accounts for, from 0 to 1. In the example it is 0.988, so about 98.8 percent of the variation in marks lines up with hours of study. A high value does not prove causation.
Can I use the regression line to predict values outside my data?
Be careful. The line is reliable inside the range of x values used to fit it. Beyond that range the relationship may bend, flatten or hit a limit, so predictions like 117 marks out of 100 become meaningless.
How many data points do I need for linear regression?
Two points define a line, but meaningful conclusions need more. A common rule of thumb is at least 20 observations for a stable line, and more if the data are noisy. Five points are fine for learning but not for decisions.
Was this guide helpful?
Continue reading
View all blogsT-Test Statistic Formula With a Worked Example
The one-sample t statistic is (mean − claimed mean) ÷ (SD ÷ √n). With mean 54, claim 50, SD 10 and n = 25, t = 2.0 on 24 degrees of freedom.
5 min read
Z-Score Formula: Standard Deviations From the Mean
A z-score subtracts the mean and divides by the standard deviation: (85 − 70) ÷ 10 = 1.5. That puts the value about the 93rd percentile if scores are normal.
5 min read
Variance and Standard Deviation: n vs n−1
Population variance divides by n, sample variance by n − 1. See both on the same six numbers: 2.92 versus 3.5, and SD 1.71 versus 1.87.
6 min read




