📈
✓ Editorially reviewed by Derek Giordano, Founder & Editor · BA Business Marketing

Linear Regression Calculator

Best-Fit Line, Slope, Intercept & R²

Updated October 2026

🧮
500+ calculators, no signup required
Finance · Health · Math · Science · Business
nnng.com
Least squares finds y = a + bx with slope b = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)². For the points (1, 2), (2, 4), (3, 5), (4, 4), (5, 5), the slope is 0.6 and the intercept 2.2, so y = 2.2 + 0.6x, with R² of 0.6. Check a residual plot; a straight line can fit curved data poorly.

What Linear Regression Tells You

Linear regression finds the straight line that best fits a set of data points — the line that comes closest, on average, to all of them. Enter your X and Y values and this calculator returns the equation of that best-fit line (y = mx + b), along with the correlation coefficient and R², which tell you how well the line actually describes your data. It's the workhorse of data analysis, used everywhere from forecasting sales to charting scientific measurements.

The method behind it is called least squares: of all possible lines, it picks the one that minimizes the total squared vertical distance between the line and the points. That single criterion produces one unique line for any dataset, which is why it's the standard.

Reading Slope, Intercept, and R²

The slope is the heart of the result: it's how much Y changes for each one-unit increase in X. A slope of 0.6 means Y rises by 0.6 every time X goes up by 1. The intercept is where the line crosses the Y axis — the predicted Y when X is zero, which sometimes is meaningful and sometimes is just a mathematical anchor.

The number people most often misread is R². It ranges from 0 to 1 and tells you the proportion of the variation in Y that the line explains. An R² of 0.9 means the line accounts for 90% of the variability — a strong fit. An R² near 0 means the line explains almost nothing, and a straight line may be the wrong model. The statistics calculator can break down the underlying spread of your data.

Correlation Is Not Causation

A strong correlation coefficient (r close to +1 or −1) means the points fall neatly along a line, but it never proves that X causes Y. Two variables can move together because both are driven by a third, hidden factor, or by pure coincidence. Regression describes a relationship; establishing cause needs more than a tidy line. Keep that distinction front of mind whenever a high R² tempts a strong conclusion.

When a Straight Line Isn't Enough

Linear regression assumes the relationship is, well, linear. If your data curves, bends, or plateaus, a straight line will fit poorly and R² will be low even when a clear pattern exists — it's just not a straight-line pattern. In those cases the right move is a different model, not a forced line. A quick scatter plot of your points is the best way to check before trusting the equation.

A Fully Worked Example

Imagine you have five months of data relating advertising spend (in thousands of dollars) to sales (in thousands of units): spend X = 2, 4, 6, 8, 10 and sales Y = 4, 6, 5, 9, 11. Working the least-squares line by hand shows precisely what the calculator computes. The method needs the means first. The mean of X is (2 + 4 + 6 + 8 + 10) ÷ 5 = 30 ÷ 5 = 6. The mean of Y is (4 + 6 + 5 + 9 + 11) ÷ 5 = 35 ÷ 5 = 7. Every later step measures how each point deviates from these two averages.

The slope is the sum of the products of paired deviations, divided by the sum of squared X deviations. For each point, compute (X − 6) and (Y − 7), then multiply them: the five products are (−4)(−3) = 12, (−2)(−1) = 2, (0)(−2) = 0, (2)(2) = 4, and (4)(4) = 16, summing to 34. The squared X deviations are 16, 4, 0, 4, and 16, summing to 40. So the slope is 34 ÷ 40 = 0.85. That means each additional thousand dollars of advertising is associated with about 850 more units sold. The intercept follows from the means: b = ȳ − m·x̄ = 7 − (0.85)(6) = 7 − 5.1 = 1.9. The best-fit line is therefore y = 0.85x + 1.9.

To judge how well that line fits, compute the correlation. You already have the numerator (34) and part of the denominator (the X deviation sum of squares, 40). You also need the sum of squared Y deviations: 9, 1, 4, 4, and 16, summing to 34. The correlation is 34 ÷ √(40 × 34) = 34 ÷ √1360 ≈ 34 ÷ 36.88 ≈ 0.922. Squaring gives R² ≈ 0.85, meaning the line explains about 85% of the variation in sales — a strong but not perfect fit, which is realistic for real-world data with other influences at play. Entering these ten numbers into the calculator returns the same equation and fit statistics immediately, including the figures the statistics calculator would give for the underlying spread.

Reading Results Without Fooling Yourself

Regression is one of the easiest tools to misuse, and most of the traps are interpretive rather than mathematical. The first is treating a high R² as proof that one variable drives the other. In the example above, advertising and sales rise together, but a tidy line cannot rule out that both are responding to something else — a holiday season, a competitor's stumble, or a general economic upswing. Establishing genuine cause requires a controlled comparison, not just a fitted slope. The line tells you how two things move together; it stays silent on why.

The second trap is extrapolation — using the line far outside the range of the data you fitted. Our example covers advertising spend from 2 to 10. The equation will happily predict sales for a spend of 100, but that prediction is essentially fiction: there is no evidence the linear relationship holds at that scale, and real systems usually saturate or break down. A model is trustworthy only within the neighborhood of the data that built it. Predictions near the edges of your range deserve caution, and predictions beyond them deserve outright skepticism.

The third issue is the outlier. Least squares minimizes squared distances, which means a single far-flung point is penalized heavily and can drag the entire line toward itself, distorting both slope and intercept. Before trusting any fit, it is worth plotting the points and asking whether one unusual observation is doing most of the work. If removing a single point changes the slope dramatically, the relationship is fragile and the R² is misleading. A clean scatter plot is the single best diagnostic, and it takes only a moment — far less time than acting on a conclusion the data does not actually support. For probability-based confidence in an estimate, the confidence interval calculator is the natural next step.

What is the least-squares regression line?
It’s the straight line that minimizes the total squared vertical distance between the line and all your data points. Of every possible line, least squares picks the single one that fits best by that measure, producing a unique equation y = mx + b for any dataset. It’s the standard method for fitting a line to data.
What does R-squared mean?
R² is the proportion of the variation in Y that the line explains, ranging from 0 to 1. An R² of 0.9 means the line accounts for 90% of the variability — a strong fit. A value near 0 means the line explains almost nothing, suggesting a straight line may be the wrong model for your data.
What is the difference between correlation and R-squared?
The correlation coefficient (r) ranges from −1 to +1 and shows the direction and strength of a linear relationship. R² is simply r squared, and it expresses the share of variance explained as a value from 0 to 1. Correlation tells you direction; R² tells you explanatory power.
Does a strong correlation prove causation?
No. A high correlation means the points fall neatly along a line, but it never proves that X causes Y. The two variables might both be driven by a hidden third factor, or the link could be coincidental. Regression describes a relationship — proving cause requires controlled evidence beyond a tidy line.
How many data points do I need?
You need at least two points to define a line, but two points always fit perfectly and tell you nothing about reliability. More points give a more trustworthy fit and a more meaningful R². As a rule of thumb, the more data you have, the more confident you can be that the line reflects a real pattern.
What if my data isn’t linear?
Linear regression assumes a straight-line relationship. If your data curves or plateaus, the line will fit poorly and R² will be low even when a clear pattern exists — it just isn’t a straight one. Plot your points first; if they bend, a different (non-linear) model is the right choice.

How to Use This Calculator

  1. Enter your X values — Type your independent-variable values separated by commas — for example, 1, 2, 3, 4, 5.
  2. Enter your Y values — Type the matching dependent-variable values in the same order, separated by commas.
  3. Read the best-fit equation — The calculator returns the line y = mx + b, computed by least squares, as the headline result.
  4. Check the slope, intercept, and R² — Use the slope to interpret the trend, the intercept for the baseline, and R² to judge how well the line fits.

Tips and Best Practices

Make sure your X and Y lists have the same number of values, entered in matching order — mismatched pairs produce a meaningless fit.

Always sketch a scatter plot before trusting the line. If the points curve, a straight-line model is the wrong tool no matter what the slope says.

Read R² as "share of variation explained." High is good, but a high R² still doesn’t prove one variable causes the other.

Watch for outliers — a single stray point can pull the whole line toward it and distort both the slope and R².

📚 Sources & References
  1. [1] NIST/SEMATECH ""e-Handbook of Statistical Methods — Linear Least Squares Regression."" NIST. itl.nist.gov
  2. [2] Khan Academy ""Least-squares regression."" Khan Academy. khanacademy.org
  3. [3] Weisstein, Eric W. ""Least Squares Fitting."" Wolfram MathWorld. mathworld.wolfram.com
  4. [4] U.S. Bureau of Labor Statistics ""Regression and statistical methods."" BLS. bls.gov
✅ Editorial Standards — Every calculator is built from peer-reviewed formulas and official data sources, editorially reviewed for accuracy, and updated regularly. Read our full methodology · About the author