Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions src/content/exercises/module-03.ts
Original file line number Diff line number Diff line change
Expand Up @@ -394,7 +394,7 @@ export const module03: ExerciseDef[] = [
{
id: 'm3-4-b',
prompt:
'How consistently do the five items measure one thing? Compute Cronbach\'s alpha for q1, q2, q3_r, q4 and q5 and store it in alpha. The formula is in the starter code.',
'How consistently do the five items agree with each other? Compute Cronbach\'s alpha for q1, q2, q3_r, q4 and q5 and store it in alpha. The formula is in the starter code.',
starterCode:
'# survey, with q3_r already added, is in your environment.\nitems <- survey[, c("q1", "q2", "q3_r", "q4", "q5")]\nk <- ncol(items)\n\n# alpha = k / (k - 1) * (1 - sum of the item variances / variance of the total)\nalpha <- ',
setupCode: REVERSED,
Expand Down Expand Up @@ -434,7 +434,7 @@ export const module03: ExerciseDef[] = [
} else if (abs(got - target) > 1e-6) {
list(pass = FALSE, message = paste0("alpha is ", round(got, 3), ", but it should be ", round(target, 3), ". The formula uses variances, var(), both for the items and for the total, rowSums()."))
} else {
list(pass = TRUE, message = paste0("Correct: alpha = ", format(round(target, 2), nsmall = 2), ". Acceptable, and the lesson shows which item is holding it back."))
list(pass = TRUE, message = paste0("Correct: alpha = ", sub("^0", "", format(round(target, 2), nsmall = 2)), ". Acceptable, and the lesson shows which item is holding it back."))
}
}
`,
Expand Down
2 changes: 1 addition & 1 deletion src/content/exercises/module-10.ts
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,7 @@ export const module10: ExerciseDef[] = [
} else if (!isTRUE(all.equal(r_w, exp_w, tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("r_workload is ", round(r_w, 3), ", but cor(d$workload, d$wellbeing) is ", round(exp_w, 3), ". Check that you correlated workload with wellbeing, and not with autonomy."))
} else {
list(pass = TRUE, message = paste0("Autonomy r = ", sub("0.", ".", round(exp_a, 3), fixed = TRUE), "; workload r = ", sub("0.", ".", round(exp_w, 3), fixed = TRUE), ". Same outcome, opposite directions. A correlation is only ever a description of these 480 employees - it is not evidence that giving someone autonomy would raise their wellbeing."))
list(pass = TRUE, message = paste0("Autonomy r = ", sub("0.", ".", format(round(exp_a, 2), nsmall = 2), fixed = TRUE), "; workload r = ", sub("0.", ".", format(round(exp_w, 2), nsmall = 2), fixed = TRUE), ". Same outcome, opposite directions. A correlation is only ever a description of these 480 employees - it is not evidence that giving someone autonomy would raise their wellbeing."))
}
}
`,
Expand Down
4 changes: 2 additions & 2 deletions src/content/exercises/module-11.ts
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ export const module11: ExerciseDef[] = [
alone <- as.vector(coef(lm(wellbeing ~ workload, data = d))["workload"])
terms_in <- names(coef(model2))
if (!identical(sort(terms_in), sort(c("(Intercept)", "autonomy", "workload")))) {
list(pass = FALSE, message = paste0("model2 should have exactly two predictors, autonomy and workload. Yours has: ", paste(setdiff(terms_in, "(Intercept)"), collapse = ", "), ". A star between predictors adds an interaction term as well; a plus sign is what you want here."))
list(pass = FALSE, message = paste0("model2 should have exactly two predictors, autonomy and workload. Yours has: ", paste(setdiff(terms_in, "(Intercept)"), collapse = ", "), if (any(grepl(":", terms_in))) ". A star between predictors adds an interaction term as well; a plus sign is what you want here." else ". Use exactly those two, joined with a plus sign."))
} else if (!is.numeric(b) || length(b) != 1L) {
list(pass = FALSE, message = "b_workload should be a single number: filter tidy() to the workload row and pull(estimate).")
} else if (isTRUE(all.equal(b, exp_autonomy, tolerance = 1e-6, check.attributes = FALSE))) {
Expand Down Expand Up @@ -248,7 +248,7 @@ export const module11: ExerciseDef[] = [
} else if (!isTRUE(all.equal(df_resid, exp_den, tolerance = 1e-9, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("df_resid is ", df_resid, " but should be ", exp_den, "."))
} else {
list(pass = TRUE, message = paste0("F(", exp_num, ", ", exp_den, ") = ", round(exp_f, 2), ", R-squared = ", sub("0.", ".", round(exp_r2, 3), fixed = TRUE), ". Those are the three numbers the first sentence of an APA regression report needs; the coefficients go in the sentences after it."))
list(pass = TRUE, message = paste0("F(", exp_num, ", ", exp_den, ") = ", round(exp_f, 2), ", R-squared = ", sub("0.", ".", round(exp_r2, 3), fixed = TRUE), ". Together with the model's p-value, those numbers make the first sentence of an APA regression report; the coefficients go in the sentences after it."))
}
}
`,
Expand Down
4 changes: 2 additions & 2 deletions src/content/exercises/module-12.ts
Original file line number Diff line number Diff line change
Expand Up @@ -180,11 +180,11 @@ export const module12: ExerciseDef[] = [
} else if (isTRUE(all.equal(b, intercept, tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("That is the intercept, ", round(intercept, 2), ". With one factor and nothing else in the model the intercept is the mean of the reference department (", exp_ref, "), not of Support and not of everybody."))
} else if (isTRUE(all.equal(b, as.vector(means[["Support"]]), tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("That is Support's own mean (", round(means[["Support"]], 2), "). The coefficient is a difference: Support's mean minus the reference department's, which is ", round(exp_b, 2), "."))
list(pass = FALSE, message = paste0("That is Support's own mean (", round(means[["Support"]], 2), "). The coefficient is a difference: Support's mean minus the reference department's, which is ", sprintf("%.2f", exp_b), "."))
} else if (!isTRUE(all.equal(b, exp_b, tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("b_support is ", round(b, 4), " but the departmentSupport coefficient is ", round(exp_b, 4), "."))
} else {
list(pass = TRUE, message = paste0("The reference is ", exp_ref, ", whose mean is the intercept, ", round(intercept, 2), ". b = ", round(exp_b, 2), " for Support means its mean is ", round(means[["Support"]], 2), ". Every coefficient in this model is a comparison with ", exp_ref, " - so none of them compares Marketing with Sales, which is what Lesson 12-3 is for."))
list(pass = TRUE, message = paste0("The reference is ", exp_ref, ", whose mean is the intercept, ", round(intercept, 2), ". b = ", sprintf("%.2f", exp_b), " for Support means its mean is ", round(means[["Support"]], 2), ". Every coefficient in this model is a comparison with ", exp_ref, " - so none of them compares Marketing with Sales, which is what Lesson 12-3 is for."))
}
}
`,
Expand Down
7 changes: 5 additions & 2 deletions src/content/exercises/module-13.ts
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,7 @@ export const module13: ExerciseDef[] = [
isTRUE(all.equal(sort(as.vector(value)), sort(as.vector(cells)), tolerance = 1e-6, check.attributes = FALSE))) found <- TRUE
}
if (!found) {
list(pass = FALSE, message = paste0("No column of cell_means holds the four mean change scores, which are ", paste(round(sort(as.vector(cells)), 2), collapse = ", "), ". Check that you summarised change (engagement_t2 minus engagement_t1) rather than engagement itself."))
list(pass = FALSE, message = paste0("No column of cell_means holds the four mean change scores, which are ", paste(sprintf("%.2f", sort(as.vector(cells))), collapse = ", "), ". Check that you summarised change (engagement_t2 minus engagement_t1) rather than engagement itself."))
} else if (!is.numeric(boost) || length(boost) != 1L) {
list(pass = FALSE, message = "boost should be a single number.")
} else if (isTRUE(all.equal(boost, corner, tolerance = 1e-6, check.attributes = FALSE))) {
Expand Down Expand Up @@ -189,7 +189,7 @@ export const module13: ExerciseDef[] = [
} else if (isTRUE(all.equal(f_training, exp_int, tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("That is the interaction's F (", round(exp_int, 2), "), on the training:mentoring row. The training main effect is on the row called training."))
} else if (isTRUE(all.equal(f_training, as.vector(default3["training", "F value"]), tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("Mentoring is still on R's default treatment contrasts, so that F (", round(as.vector(default3["training", "F value"]), 2), ") is not the main effect of training. With mentoring treatment-coded, a type III test of training asks about training among employees with NO mentoring only: a simple effect wearing a main effect's name. With contrasts = list(training = contr.sum, mentoring = contr.sum) the same row becomes ", round(exp_f, 2), "."))
list(pass = FALSE, message = paste0("Mentoring is on R's default treatment contrasts, so that F (", round(as.vector(default3["training", "F value"]), 2), ") is not the main effect of training. With mentoring treatment-coded, a type III test of training asks about training among employees with NO mentoring only: a simple effect wearing a main effect's name. With contrasts = list(training = contr.sum, mentoring = contr.sum) the same row becomes ", round(exp_f, 2), "."))
} else if (isTRUE(all.equal(f_training, as.vector(type2["training", "F value"]), tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("That is the type II F (", round(as.vector(type2["training", "F value"]), 2), "). Type II tests each main effect while ignoring the interaction, which is only defensible when the interaction is negligible. This exercise asks for type III."))
} else if (!isTRUE(all.equal(f_training, exp_f, tolerance = 1e-6, check.attributes = FALSE))) {
Expand Down Expand Up @@ -242,12 +242,15 @@ export const module13: ExerciseDef[] = [
exp_no <- as.vector(cells["Yes", "No"] - cells["No", "No"])
exp_yes <- as.vector(cells["Yes", "Yes"] - cells["No", "Yes"])
overall <- mean(d$change[d$training == "Yes"]) - mean(d$change[d$training == "No"])
t2cells <- tapply(d$engagement_t2, list(d$training, d$mentoring), mean)
if (!is.numeric(no_m) || length(no_m) != 1L || !is.numeric(yes_m) || length(yes_m) != 1L) {
list(pass = FALSE, message = "Both should be single numbers.")
} else if (isTRUE(all.equal(no_m, yes_m, tolerance = 1e-12, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("Your two simple effects are identical, so you have reported the overall training effect (", round(overall, 2), ") twice. The whole point of a simple effect is that it differs between the levels of the other factor - here they are ", round(exp_no, 2), " and ", round(exp_yes, 2), "."))
} else if (isTRUE(all.equal(no_m, -exp_no, tolerance = 1e-6, check.attributes = FALSE)) && isTRUE(all.equal(yes_m, -exp_yes, tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = "Both of your differences run backwards. Each simple effect is the training-Yes cell minus the training-No cell within that level of mentoring.")
} else if (isTRUE(all.equal(no_m, as.vector(t2cells["Yes", "No"] - t2cells["No", "No"]), tolerance = 1e-6, check.attributes = FALSE)) && isTRUE(all.equal(yes_m, as.vector(t2cells["Yes", "Yes"] - t2cells["No", "Yes"]), tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = "Those are differences in engagement at time 2, not in the change score. Compute change = engagement_t2 - engagement_t1 first, then take the differences.")
} else if (!isTRUE(all.equal(no_m, exp_no, tolerance = 1e-6, check.attributes = FALSE))) {
list(pass = FALSE, message = paste0("effect_no_mentoring is ", round(no_m, 4), " but the training effect among employees without mentoring is ", round(exp_no, 4), ". Hold mentoring at No and take the difference across training."))
} else if (!isTRUE(all.equal(yes_m, exp_yes, tolerance = 1e-6, check.attributes = FALSE))) {
Expand Down
2 changes: 1 addition & 1 deletion src/content/exercises/module-15.ts
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,7 @@ export const module15: ExerciseDef[] = [
}
}
if (!found) {
list(pass = FALSE, message = paste0("No column of rate_by_third holds the three leaving rates, which are ", paste(round(exp_rates, 3), collapse = ", "), " from the lowest third of wellbeing to the highest. Split on wellbeing, not on left_company."))
list(pass = FALSE, message = paste0("No column of rate_by_third holds the three leaving rates, which are ", paste(round(exp_rates_ntile, 3), collapse = ", "), " from the lowest third of wellbeing to the highest. Split on wellbeing, not on left_company."))
} else {
list(pass = TRUE, message = paste0(exp_n, " of the ", nrow(d), " fitted values are impossible probabilities. And the descriptives say the effect is real: ", round(100 * shown[1], 1), "% of the least happy third left, against ", round(100 * shown[3], 1), "% of the happiest. A model that predicts a negative probability for the very employees it should be most confident about is the wrong shape, not the wrong data."))
}
Expand Down
2 changes: 2 additions & 0 deletions src/content/exercises/module-16.ts
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,8 @@ export const module16: ExerciseDef[] = [
list(pass = FALSE, message = "That is a 90% interval. A 95% interval leaves 2.5% in each tail: qbeta(c(0.025, 0.975), ...).")
} else if (close(ci, binom.test(k, n)$conf.int)) {
list(pass = FALSE, message = "That is the frequentist confidence interval from binom.test(). Here the interval should come from the posterior, with qbeta().")
} else if (close(ci, qbeta(c(0.025, 0.975), 1 + sum(d$left_company), 1 + nrow(d) - sum(d$left_company)))) {
list(pass = FALSE, message = "That is the interval for the whole company. Use Engineering's counts: left_eng out of n_eng.")
} else if (!close(ci, expected)) {
list(pass = FALSE, message = paste0("ci_engineering is [", paste(round(ci, 3), collapse = ", "), "]. The posterior is Beta(1 + ", k, ", 1 + ", n - k, "): leavers in the first shape, stayers in the second."))
} else {
Expand Down
2 changes: 1 addition & 1 deletion src/content/exercises/module-17.ts
Original file line number Diff line number Diff line change
Expand Up @@ -103,7 +103,7 @@ export const module17: ExerciseDef[] = [
} else if (!identical(as.character(got$call$rotation), "promax")) {
list(pass = FALSE, message = "Two factors, good. Let them correlate with rotation = \\"promax\\": pressure and support are unlikely to be unrelated.")
} else {
list(pass = TRUE, message = paste0("Correct. The chi-square test of fit gives p = ", format(round(got$PVAL, 2), nsmall = 2), ", so there is no evidence that two factors fall short. Print it with print(efa, cutoff = 0.3) to see which items load where."))
list(pass = TRUE, message = paste0("Correct. The chi-square test of fit gives p = ", sub("^0", "", format(round(got$PVAL, 2), nsmall = 2)), ", so there is no evidence that two factors fall short. Print it with print(efa, cutoff = 0.3) to see which items load where."))
}
}
`,
Expand Down
7 changes: 4 additions & 3 deletions src/content/lessons/00-1-rstudio-and-projects.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -85,9 +85,10 @@ top right of the window shows the project name, the Files pane shows the new
folder, and the folder contains a file ending in `.Rproj`.

To come back to the project later, double-click that `.Rproj` file, or pick the
project from **File > Recent Projects**. Opening RStudio on its own and then
opening a script does not open the project, and the working directory will be
wrong.
project from **File > Recent Projects**. Opening RStudio on its own reopens
whichever project was open last, which may not be this one, and opening a script
does not switch projects. If the top right does not show this project's name,
the working directory will be wrong.

## A folder layout that works

Expand Down
3 changes: 2 additions & 1 deletion src/content/lessons/01-2-functions-and-help.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -86,7 +86,8 @@ inside out.
length(seq(from = 0, to = 100, by = 10))`} />

Read the first line inside out: average the scores, ignoring the missing one,
then round to one decimal. It is compact, and once there are four or five
then round to one decimal. R prints 64.2, not the 64.3 you may have learned:
when a number sits exactly halfway, R rounds to the even digit. It is compact, and once there are four or five
functions stacked up it becomes unreadable. The pipe you met in Module 0 is
R's answer to exactly that problem.

Expand Down
2 changes: 1 addition & 1 deletion src/content/lessons/02-2-pipe-and-verbs.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ A minus sign drops instead of keeps: `select(-site)` returns everything except
head(4)`} />

`mutate()` works on whole columns at once, so every row gets its own result and
the new column is as long as the table. Nothing has been saved yet: every block on this page has printed a new
the new column is as long as the table. Nothing has been saved yet: every pipeline on this page has printed a new
table and thrown it away. Storing it takes the arrow.

<Exercise id="m2-2-a" />
Expand Down
6 changes: 3 additions & 3 deletions src/content/lessons/03-4-scales-and-reliability.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -85,7 +85,7 @@ total_variance <- var(rowSums(items))

k / (k - 1) * (1 - item_variances / total_variance)`} />

Alpha is 0.74. Rough conventions call .70 acceptable and .80 good for research
Alpha is .74. Rough conventions call .70 acceptable and .80 good for research
use. Alpha also rises with the number of items, so a long scale can reach .80
with items that agree only weakly.

Expand All @@ -104,7 +104,7 @@ same thing.

sapply(names(items), function(item) alpha_of(items[, names(items) != item]))`} />

Dropping any of q1 to q4 lowers alpha. Dropping q5 raises it to 0.82, and its
Dropping any of q1 to q4 lowers alpha. Dropping q5 raises it to .82, and its
correlations with the other items were close to zero. The coffee question does
not belong in a wellbeing scale. With a real questionnaire, you would look at
the wording before you drop anything, and you would report which items you
Expand All @@ -118,7 +118,7 @@ you when an item looks as if it still needs reversing.
question="You forget to reverse q3 and compute alpha on the original five items. What happens?"
choices={[
{ text: 'Nothing much, alpha only looks at variances', response: 'The total variance depends on how the items move together, and an unreversed item moves against the rest.' },
{ text: 'Alpha collapses towards zero, because q3 cancels the other items out', correct: true, response: 'Right. Here it drops from 0.74 to 0.02. A suspiciously low alpha is often an item nobody reversed.' },
{ text: 'Alpha collapses towards zero, because q3 cancels the other items out', correct: true, response: 'Right. Here it drops from .74 to .02. A suspiciously low alpha is often an item nobody reversed.' },
{ text: 'Alpha goes up, because there is more variation', response: 'The items vary as much as before, but the total varies less, because q3 pulls against the others. Alpha goes down.' },
]}
/>
Expand Down
6 changes: 4 additions & 2 deletions src/content/lessons/04-2-boxplots-and-facets.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -21,8 +21,10 @@ computed:
- the box spans the **quartiles** — the middle half of the department, so its
height is the IQR;
- the whiskers reach out to the furthest observation within 1.5 IQRs of the box
(ggplot2 measures from the quartiles; the textbook describes the distance from
the median, so its whiskers can end slightly differently);
(ggplot2 measures from the quartiles, as does the textbook's worked wage
example; the textbook's prose measures from the median instead, which gives
shorter whiskers and flags more points: here Sales would show seven points
below its whisker instead of none);
- anything beyond that is drawn as a point, and is worth looking at rather than
deleting.

Expand Down
4 changes: 3 additions & 1 deletion src/content/lessons/05-1-density-and-area.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,9 @@ population <- read.csv("data/wellbeing-population.csv", stringsAsFactors = TRUE)

population %>% summarise(mu = mean(exam_score), sigma = sd(exam_score), n = n())`} />

Two numbers: a mean of about 73.3 and a standard deviation of about 12.4. The
Two numbers: a mean of about 73.3 and a standard deviation of about 12.4. (sd()
divides by n − 1; the textbook's population σ divides by N and gives 12.35
instead of 12.36, a difference that changes nothing in this module.) The
claim of this lesson is that those two numbers are almost the whole story.

<CodeBlock id="hist" code={`library(ggplot2)
Expand Down
Loading