Create section for linear regression chapter consisting of:
Simulate Y = a+bX+cX^2 + U where X is uniformly distributed (say sample size 500) and U is shifted exponentially distributed with parameter 1. So just consider an exponentially distributed random variable with parameter 1 and subtract 1 from the result. Then run the simple linear regression model on (X,Y) and plot the regression line in a scatterplot of X vs Y, and the summary. Make note of p-values and R^2 in the summary. Then we do the residual analysis. Plot x against residuals and also plot a histogram of the residuals. Note that the histogram shows errors are not normal. This means p-values and confidence intervals for the coefficients are only an approximation (a good one in this case since sample size is high). Then note the residuals plot shows non-linear. In particular it shows quadratic dependence. Then run the model again with the extra X^2 term included. Now plot the regression parabola in the scatterplot and the residuals again to note the model is improved also mention the R^2 improving. Invite students to try the simulation again with other non-linear behaviour included line third order or sin(X). Next we do another simulation to consider hetero-skedasticity: simulate Y = a + bX + U as before but no U is normally distributed with mean 0 and standard deviation X. Again plot scatter plot against estimated regression line and summary and note relevant values. Also histogram of residuals (should show normal now) and scatterplot of x against residuals, which should show hetero-skedasticity. Then apply transformations to Y (say try both ln(Y) and sqrt(Y)) and run the model again to note that the hetero-skedasticity is diminished.
Create section for linear regression chapter consisting of:
Simulate Y = a+bX+cX^2 + U where X is uniformly distributed (say sample size 500) and U is shifted exponentially distributed with parameter 1. So just consider an exponentially distributed random variable with parameter 1 and subtract 1 from the result. Then run the simple linear regression model on (X,Y) and plot the regression line in a scatterplot of X vs Y, and the summary. Make note of p-values and R^2 in the summary. Then we do the residual analysis. Plot x against residuals and also plot a histogram of the residuals. Note that the histogram shows errors are not normal. This means p-values and confidence intervals for the coefficients are only an approximation (a good one in this case since sample size is high). Then note the residuals plot shows non-linear. In particular it shows quadratic dependence. Then run the model again with the extra X^2 term included. Now plot the regression parabola in the scatterplot and the residuals again to note the model is improved also mention the R^2 improving. Invite students to try the simulation again with other non-linear behaviour included line third order or sin(X). Next we do another simulation to consider hetero-skedasticity: simulate Y = a + bX + U as before but no U is normally distributed with mean 0 and standard deviation X. Again plot scatter plot against estimated regression line and summary and note relevant values. Also histogram of residuals (should show normal now) and scatterplot of x against residuals, which should show hetero-skedasticity. Then apply transformations to Y (say try both ln(Y) and sqrt(Y)) and run the model again to note that the hetero-skedasticity is diminished.