Regression analysis
Fit a simple or multiple linear regression by least squares and get the coefficients, intercept, R², adjusted R², residual standard error, per-coefficient standard errors, residuals and a ready-to-use prediction function.
Runs in your browserEvery computation happens in your browser — your data never leaves this device.
Input
Least squares solves the normal equations (XᵀX)β = Xᵀy with column-pivoted Gaussian elimination; the coefficient standard errors come from the diagonal of σ²(XᵀX)⁻¹.
Regression result
y = 2.2000 + 0.6000 × x12.20000.93810.60000.28280.60000.46670.894453-0.8000, 0.6000, 1.0000, -0.6000, -0.2000
Prediction
Leave empty to skip; the length must match the number of columns in X.
What this tool does
- Fit a trend line by least squares: paste paired data such as ad spend against revenue and get the slope, the intercept and an equation you can predict with.
- Multiple regression: control for several drivers (floor area, building age, floor level) at once and read each coefficient together with its standard error to see what really matters.
- Check the fit quality: read R² next to adjusted R² and scan the residual list — residuals that grow with x are the classic sign of a missing curve.
- Homework and lab data: get coefficients, standard errors, degrees of freedom and residuals in one pass for a results table, plus a prediction function for follow-up values.
Example
Input
X (one sample per line) 1 / 2 / 3 / 4 / 5; Y 2, 4, 5, 4, 5
Output
Equation y = 2.2000 + 0.6000 × x1, intercept 2.2000 (standard error 0.9381), x1 coefficient 0.6000 (standard error 0.2828), R² = 0.6000, adjusted R² = 0.4667, residual standard error 0.8944, 3 degrees of freedom, residuals -0.8000, 0.6000, 1.0000, -0.6000, -0.2000
Enter 10 as the prediction sample and you get 2.2000 + 0.6000 × 10 = 8.2000. For multiple regression write several columns per line, e.g. "1, 5" means x1 = 1 and x2 = 5.
Frequently asked questions
How many samples do I need?
The tool requires strictly more samples than predictors plus one, because the degrees of freedom n − k − 1 must stay positive or the residual standard error and coefficient standard errors are undefined. In practice, budget about ten samples per predictor, otherwise R² is inflated and the coefficients unstable.
R² versus adjusted R²?
R² can only go up when you add a predictor, even pure noise. Adjusted R² penalises for degrees of freedom: 1 − (1 − R²)(n − 1)/(n − k − 1). The more predictors and the fewer samples, the heavier the penalty, so adjusted R² is the fairer comparison across models of different size.
Why do I get "the design matrix is rank deficient"?
Some predictors are perfectly collinear — two identical columns, or one column that is a linear combination of the others — so (XᵀX) is not invertible and the coefficients are not unique. Drop the redundant column, or switch to a regularised method such as ridge regression.
What is the standard error for?
A coefficient standard error measures how precisely that coefficient is estimated: coefficient divided by standard error is the t statistic used to test whether the predictor differs from zero. Smaller standard errors mean the estimate is more stable for the data you have.
Can I extrapolate with the fitted equation?
A linear model is only reliable inside the range of the data. Feeding a prediction sample far beyond the training values is extrapolation and the error grows quickly. Remember too that a regression describes association, not causation.
Keywords:linear regressionmultiple regressionleast squaresr squaredadjusted r squaredresiduals线性回归多元回归最小二乘回归系数残差预测