What regression does
Regression fits a line (or a plane, with more variables) through data to describe and predict a numerical outcome. The outcome is the dependent variable, usually called y. The drivers are independent variables, called x. A simple regression has one x, such as advertising spend. A multiple regression has several, such as advertising spend, price and season.
The simple model is y = b0 + b1 x. The intercept b0 is the predicted y when x is zero. The slope b1 is the predicted change in y for a one-unit increase in x. Least squares chooses the line that makes the sum of squared vertical gaps between the data and the line as small as possible.
Worked example: calculating a simple regression by hand
A business records monthly advertising spend and sales, both in thousands of dollars, for six months.
| Month | Advertising x | Sales y |
|---|---|---|
| 1 | 2 | 20 |
| 2 | 3 | 24 |
| 3 | 4 | 29 |
| 4 | 5 | 31 |
| 5 | 6 | 38 |
| 6 | 7 | 40 |
Step 1, means. Mean of x is 4.5. Mean of y is 182 / 6 = 30.33.
Step 2, sums of squares and products. The sum of (x minus mean x) squared is 6.25 + 2.25 + 0.25 + 0.25 + 2.25 + 6.25 = 17.5. The sum of (x minus mean x) times (y minus mean y) is 25.83 + 9.50 + 0.67 + 0.33 + 11.50 + 24.17 = 72.0.
Step 3, slope and intercept. b1 = 72.0 / 17.5 = 4.114. b0 = 30.33 - 4.114 x 4.5 = 11.82.
The fitted line is sales = 11.82 + 4.114 x advertising.
| x | Actual y | Predicted y | Residual (actual minus predicted) |
|---|---|---|---|
| 2 | 20 | 20.05 | -0.05 |
| 3 | 24 | 24.16 | -0.16 |
| 4 | 29 | 28.27 | 0.73 |
| 5 | 31 | 32.39 | -1.39 |
| 6 | 38 | 36.50 | 1.50 |
| 7 | 40 | 40.62 | -0.62 |
The residuals add to approximately zero, as they should. The sum of squared residuals (SSE) is about 5.13. The total variation in sales around its mean (SST) is 301.33. So R-squared = 1 - 5.13 / 301.33 = 0.983. The standard error of the estimate is the square root of 5.13 / 4, which is 1.13.
Interpreting the result in business terms
- Slope: each additional $1,000 of advertising is associated with about $4,114 more in sales, on average, in the range of data observed.
- Intercept: the model predicts sales of about $11,820 with no advertising. Be careful, because zero advertising is outside the observed range of 2 to 7, so this is an extrapolation and not a reliable claim.
- R-squared: advertising explains about 98 percent of the variation in sales across these six months. That is very high and unusual in real data. With only six points, treat it cautiously.
- Prediction: with advertising of 5.5, predicted sales are 11.82 + 4.114 x 5.5 = $34.45 thousand, which is inside the data range, so this is a more trustworthy interpolation.
Say associated with rather than caused by unless the data come from an experiment. Regression describes patterns, and it does not by itself prove that advertising causes sales. Seasonality or pricing could drive both.
Reading Excel output
Excel's regression tool (Data Analysis ToolPak) produces three blocks. Here is the output for the example, with explanations.
| Regression statistics | Value | Meaning |
|---|---|---|
| Multiple R | 0.991 | Correlation between actual and predicted y |
| R Square | 0.983 | Share of variation in y explained |
| Adjusted R Square | 0.979 | R-squared adjusted for the number of predictors, better for comparing models |
| Standard Error | 1.13 | Typical size of prediction errors, in y units |
| Observations | 6 | Sample size |
| Coefficients | Coefficient | Standard error | t Stat | P-value |
|---|---|---|---|---|
| Intercept | 11.82 | 1.30 | 9.07 | 0.0008 |
| Advertising | 4.114 | 0.271 | 15.2 | less than 0.001 |
The t statistic is the coefficient divided by its standard error (4.114 / 0.271 = 15.2). It tests whether the true slope is zero. A small p-value, here far below 0.05, means advertising has a statistically significant relationship with sales. The F-test in the ANOVA block tests whether the model as a whole is useful, and with one predictor it is the square of the slope's t statistic (15.2 squared is about 231). The 95 percent confidence interval for the slope is roughly 4.114 plus or minus 2.78 x 0.271, which is from about 3.36 to 4.87.
Working on this assignment now? Get a price for help with your paper.
Get an instant quoteMultiple regression
Adding variables lets you separate effects. Suppose sales depend on advertising and price. The model is sales = b0 + b1 x advertising + b2 x price. Interpretation changes slightly: b1 is the change in sales for a one-unit increase in advertising holding price constant, and b2 is the effect of price holding advertising constant.
| Feature | How to handle it |
|---|---|
| Categorical predictors (for example holiday month yes or no) | Create a dummy variable coded 0 or 1; its coefficient is the difference in y between the two groups, holding the other variables constant |
| More than two categories | Use one fewer dummy than the number of categories; the left-out category is the baseline |
| Adjusted R-squared | Use it rather than R-squared to compare models with different numbers of variables, since R-squared never falls when you add variables |
| Multicollinearity | When predictors are highly correlated with each other, coefficients become unstable; check correlations or variance inflation factors (VIF above about 5 to 10 is a warning) |
| Which variables to keep | Keep those with sensible reasoning and statistical significance, not just a high R-squared |
Assumptions and how to check them
| Assumption | What it means | How to check | If violated |
|---|---|---|---|
| Linearity | The relationship between x and y is a straight line | Scatter plot, residuals versus fitted values | Transform variables or add a curved term |
| Independence | Errors are not related to each other | Plot residuals over time; Durbin-Watson test for time series | Use time-series methods |
| Constant variance | The spread of residuals is similar across fitted values | Residual plot with a fan shape is a warning | Transform y or use robust methods |
| Normal errors | Residuals are roughly normal | Histogram or normal probability plot | Matters mostly for small samples |
In an assignment, include a scatter plot and a residual plot, and write one or two sentences on whether the assumptions look reasonable. Also check for outliers and influential points, because a single unusual observation can change the line a lot.
How to write up regression results
A short write-up (using the example)
A simple linear regression of monthly sales on advertising spend, using six months of data, found a significant positive relationship (b = 4.11, t(4) = 15.2, p < .001). Each additional $1,000 spent on advertising was associated with about $4,100 in extra sales. The model explained 98 percent of the variance in sales (R-squared = .98). Because the sample is small and covers a limited range of spending (from $2,000 to $7,000), the results should not be extrapolated beyond that range, and the data do not show that advertising causes the increase in sales.
- State the model Say what y and x are, and the number of observations.
- Report the key statistics Coefficients, significance, R-squared and sample size.
- Interpret in business terms Use units and plain language, and say what a change in x means for y.
- Check and mention assumptions Include at least one diagnostic plot.
- Avoid overclaiming Use associated with unless causation can be justified, and avoid extrapolation.
Common mistakes
- Confusing correlation with causation A strong relationship does not show that x causes y.
- Extrapolating Predicting far outside the range of the data is unreliable.
- Relying on R-squared alone A high R-squared does not mean the model is correct. Check residuals.
- Ignoring units Always interpret the slope in the units of x and y.
- Adding variables without thought Each variable should have a reason, not just improve the fit.
- Forgetting categorical variables need dummies Do not enter categories as numbers 1, 2, 3.
If you want help running regressions or writing up results, you can order a data analysis report and upload your data.