Linear regression finds the straight line that best summarizes how one variable changes with another. “Best” has a precise meaning: the line makes the sum of squared vertical distances from the points as small as possible. Paste your x and y values to get the regression equation, R², standard errors and p-values, a prediction with confidence and prediction intervals, the full residual table and a scatter plot with the fitted line.
How to use the linear regression calculator
- Enter the predictor values in X and the response values, in the same order, in Y.
- Optionally enter an x value to predict y. Leave it blank to skip the prediction.
- Choose the confidence level for the slope interval and the prediction intervals (95% is standard).
- Read the equation on the tape, check the scatter plot for curves or outliers, and use the residual table to see how far each point falls from the line.
Least-squares formulas
With Sxx = Σ(x − x̄)² and Sxy = Σ(x − x̄)(y − ȳ):
Fit quality and uncertainty come from the residuals e = y − ŷ:
A prediction at x₀ has the interval ŷ ± t* × s × √(1 + 1/n + (x₀ − x̄)²/Sxx); dropping the “1 +” gives the narrower interval for the mean response.
Worked example
A small business records monthly advertising spend (x, in thousands of dollars) and sales (y, in thousands):
| x | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| y | 3.5 | 4.0 | 7.2 | 6.1 | 9.8 | 8.7 | 12.9 | 11.8 |
- Means: x̄ = 4.5 and ȳ = 8.0.
- Sxx = 42 and Sxy = 55.4.
- Slope: 55.4 ÷ 42 = 1.3190. Intercept: 8.0 − 1.3190 × 4.5 = 2.0643.
- Equation: ŷ = 2.0643 + 1.3190x. Each extra $1,000 of advertising goes with about $1,319 more in sales.
- SSE = 9.6048 against a total of 82.68, so R² = 0.8838.
- The slope’s standard error is 0.1952, so t = 6.756 with 6 df and p = 0.0005. The 95% interval for the slope is 0.841 to 1.797.
- At x = 6.5, the predicted sales are 10.64. The 95% interval for the average month at that spend is 9.19 to 12.09, while a single future month could plausibly land anywhere from 7.22 to 14.06.
Checking the fit
Look at the residuals
A good linear fit leaves residuals that scatter randomly around zero. Patterns signal problems: a curve (residuals positive at both ends, negative in the middle) means the relationship is not linear; a funnel shape means the spread grows with x. Residuals always sum to zero in least squares, which the table confirms.
R² is not everything
R² can be high for a curved relationship and low for a correct model with noisy data. It also always rises when data cover a wider x range. Judge a model by the residual pattern, the size of s in the units of y, and whether the slope makes sense.
Influential points
A point far out on the x axis can pull the line toward itself. Try removing it and refitting; if the slope changes a lot, report both fits. The outlier calculator helps flag extreme values first.
Regression and correlation
In simple regression, R² equals the square of Pearson’s r, and the slope’s t-test is identical to the test of r = 0. For the correlation alone, use the correlation coefficient calculator. To smooth a time series instead of fitting a line, try the moving average calculator.
Frequently asked questions
What does the slope of a regression line mean?
It is the predicted change in y for a one-unit increase in x. A slope of 1.319 means that each extra unit of x is associated with about 1.32 more units of y on average. Its sign matches the sign of the correlation.
What is a good R² value?
There is no universal cutoff. Controlled lab measurements often reach 0.95 or higher, while models of human behavior may be useful at 0.2. R² tells you the share of variation explained, not whether the model is correct; a curved pattern can still give a high R² with a straight line.
What is the difference between a confidence interval and a prediction interval?
The confidence interval brackets the average y for all cases with that x value. The prediction interval brackets a single new observation, so it also includes the scatter of individual points around the line and is always wider.
Can I use the line to predict outside my data?
You can compute a value, but it is extrapolation: it assumes the straight-line pattern continues where you have no evidence. Predictions far outside the observed x range are often badly wrong. The calculator warns you when x is outside the range.
Does it matter which variable is x and which is y?
Yes. Least squares minimizes vertical distances, so regressing y on x gives a different line from regressing x on y unless the correlation is perfect. Put the variable you want to predict in Y and the predictor in X.