Best-Fit Line, Slope, Intercept & R²
Updated October 2026
Linear regression finds the straight line that best fits a set of data points — the line that comes closest, on average, to all of them. Enter your X and Y values and this calculator returns the equation of that best-fit line (y = mx + b), along with the correlation coefficient and R², which tell you how well the line actually describes your data. It's the workhorse of data analysis, used everywhere from forecasting sales to charting scientific measurements.
The method behind it is called least squares: of all possible lines, it picks the one that minimizes the total squared vertical distance between the line and the points. That single criterion produces one unique line for any dataset, which is why it's the standard.
The slope is the heart of the result: it's how much Y changes for each one-unit increase in X. A slope of 0.6 means Y rises by 0.6 every time X goes up by 1. The intercept is where the line crosses the Y axis — the predicted Y when X is zero, which sometimes is meaningful and sometimes is just a mathematical anchor.
The number people most often misread is R². It ranges from 0 to 1 and tells you the proportion of the variation in Y that the line explains. An R² of 0.9 means the line accounts for 90% of the variability — a strong fit. An R² near 0 means the line explains almost nothing, and a straight line may be the wrong model. The statistics calculator can break down the underlying spread of your data.
A strong correlation coefficient (r close to +1 or −1) means the points fall neatly along a line, but it never proves that X causes Y. Two variables can move together because both are driven by a third, hidden factor, or by pure coincidence. Regression describes a relationship; establishing cause needs more than a tidy line. Keep that distinction front of mind whenever a high R² tempts a strong conclusion.
Linear regression assumes the relationship is, well, linear. If your data curves, bends, or plateaus, a straight line will fit poorly and R² will be low even when a clear pattern exists — it's just not a straight-line pattern. In those cases the right move is a different model, not a forced line. A quick scatter plot of your points is the best way to check before trusting the equation.
Imagine you have five months of data relating advertising spend (in thousands of dollars) to sales (in thousands of units): spend X = 2, 4, 6, 8, 10 and sales Y = 4, 6, 5, 9, 11. Working the least-squares line by hand shows precisely what the calculator computes. The method needs the means first. The mean of X is (2 + 4 + 6 + 8 + 10) ÷ 5 = 30 ÷ 5 = 6. The mean of Y is (4 + 6 + 5 + 9 + 11) ÷ 5 = 35 ÷ 5 = 7. Every later step measures how each point deviates from these two averages.
The slope is the sum of the products of paired deviations, divided by the sum of squared X deviations. For each point, compute (X − 6) and (Y − 7), then multiply them: the five products are (−4)(−3) = 12, (−2)(−1) = 2, (0)(−2) = 0, (2)(2) = 4, and (4)(4) = 16, summing to 34. The squared X deviations are 16, 4, 0, 4, and 16, summing to 40. So the slope is 34 ÷ 40 = 0.85. That means each additional thousand dollars of advertising is associated with about 850 more units sold. The intercept follows from the means: b = ȳ − m·x̄ = 7 − (0.85)(6) = 7 − 5.1 = 1.9. The best-fit line is therefore y = 0.85x + 1.9.
To judge how well that line fits, compute the correlation. You already have the numerator (34) and part of the denominator (the X deviation sum of squares, 40). You also need the sum of squared Y deviations: 9, 1, 4, 4, and 16, summing to 34. The correlation is 34 ÷ √(40 × 34) = 34 ÷ √1360 ≈ 34 ÷ 36.88 ≈ 0.922. Squaring gives R² ≈ 0.85, meaning the line explains about 85% of the variation in sales — a strong but not perfect fit, which is realistic for real-world data with other influences at play. Entering these ten numbers into the calculator returns the same equation and fit statistics immediately, including the figures the statistics calculator would give for the underlying spread.
Regression is one of the easiest tools to misuse, and most of the traps are interpretive rather than mathematical. The first is treating a high R² as proof that one variable drives the other. In the example above, advertising and sales rise together, but a tidy line cannot rule out that both are responding to something else — a holiday season, a competitor's stumble, or a general economic upswing. Establishing genuine cause requires a controlled comparison, not just a fitted slope. The line tells you how two things move together; it stays silent on why.
The second trap is extrapolation — using the line far outside the range of the data you fitted. Our example covers advertising spend from 2 to 10. The equation will happily predict sales for a spend of 100, but that prediction is essentially fiction: there is no evidence the linear relationship holds at that scale, and real systems usually saturate or break down. A model is trustworthy only within the neighborhood of the data that built it. Predictions near the edges of your range deserve caution, and predictions beyond them deserve outright skepticism.
The third issue is the outlier. Least squares minimizes squared distances, which means a single far-flung point is penalized heavily and can drag the entire line toward itself, distorting both slope and intercept. Before trusting any fit, it is worth plotting the points and asking whether one unusual observation is doing most of the work. If removing a single point changes the slope dramatically, the relationship is fragile and the R² is misleading. A clean scatter plot is the single best diagnostic, and it takes only a moment — far less time than acting on a conclusion the data does not actually support. For probability-based confidence in an estimate, the confidence interval calculator is the natural next step.
Make sure your X and Y lists have the same number of values, entered in matching order — mismatched pairs produce a meaningless fit.
Always sketch a scatter plot before trusting the line. If the points curve, a straight-line model is the wrong tool no matter what the slope says.
Read R² as "share of variation explained." High is good, but a high R² still doesn’t prove one variable causes the other.
Watch for outliers — a single stray point can pull the whole line toward it and distort both the slope and R².