Correct the keys, help text, figures and report tables - #98
Merged
Merged
Conversation
…ll panels From the pre-1.0 audit. Values that change: - Constructs coded as numbers indexed the count table by position, so the target count, Csv, p, the decision and the handoff items could be wrong. Labels are now compared as text. - An item sorted by too few judges for any count to reach alpha is "Insufficient panel" (status "Insufficient data"), not "Review". It stays out of scale means, and the handoff rule no longer reads ">= NA of 4". - csv_binom_test() builds its interval at 1 - alpha, not always 95%. - sort_power() gives a power of 0, not NA, when no count can reach alpha. - A scale mean on a published Colquitt band minimum falls in that band. - Leading and trailing spaces in labels are ignored. - Items keep the order of the data (or of factor levels), not text order. Other fixes: - Factor columns with different level sets no longer stop the sort. - Alpha prints with the digits given (.001, not .00), in every workflow and in the handoff rule. - csv_binom_test() prints its verdict first, without "n.s..". - p0 is described as a null probability, not a chance rate, with Howard and Melloy's caution that .5 is arbitrary and lenient. - Vignettes: triangles, not diamonds. README: names the walkthrough files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…umns From the pre-1.0 audit and its completeness check. Values that change: - The default relevance cut was hi - 1, the bottom of a 0/1 scale, so every rating counted and an unendorsed item got I-CVI 1.00. The default is now hi on a two-point scale, and a cut at the bottom of any scale is refused, in expert_validity(), judge_validity() and delphi_validity(). - aikens_v() requires lo and hi. It assumed 1-5 while the workflows assume 1-4, so aikens_v(R) on 1-4 ratings gave .75 for an item rated 4 by all. - A column whose name looks like a rater ID stops every judge-by-item function, where it was analyzed as an item. - An essentiality item rated by too few experts for the exact test is "Insufficient panel", not "Review", and the handoff rule has no "NA". - Named essential counts keep their names, into the handoff. - Lawshe's table at nine panelists: 8 of 9 meets his .78, which is 8 of 9 rounded. The comparison uses the count each tabled value implies. - A seed no longer leaves the session's random stream reset. - The AC1 bootstrap scores every resample on the categories of the full data. Other fixes: - The relevance print states the scale and the cut. - An out-of-range rating names its column and the scale. - cvr() and panel_agreement() no longer print NA for an undefined value. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Four reviewers checked the first commit. Two regressions it introduced are fixed, and the rest tightened: - A NaN in a numeric assigned column stays missing; converting to text had turned it into a judge's assignment. - Numeric labels are written without scientific notation, one value at a time, so the integer 100000 and the double 1e5 are one label. - Non-breaking spaces are trimmed too, in item, judge and construct labels and in the names that key orbiting_r. - Scale means are back to averaging every item with a Psa and a Csv. Excluding small-panel items made the means depend on alpha. - summary() counts "Too few judges" apart from "Insufficient data". - The legacy comparison keeps each row on one line and says so when no item has a decision to compare. - .fmt_alpha() no longer depends on the digits or OutDec options, and shows a computed level to three significant digits. - Wording: singular for one item, "1 of 1 judge", and no claim that the interval excludes p0 when p equals alpha exactly. - NEWS names every changed output, widens the rerun advice to factor-coded constructs, and files the cross-workflow changes under their own heading. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From the pre-1.0 audit. Values that change: - Text round labels are ordered by the number each carries, not by the order of the rows; labels with no number stop and ask for a factor. Row order had decided silently, giving the wrong last round and consensus decisions. - compare_rounds() treats a changed panel size as a changed decision rule, ignores the seed and resample count, refuses fits from different expert-panel modes, and refuses labels that overwrite a column. - The Delphi handoff rule gives the round by position (no "round 4 of 3"), says the threshold was supplied rather than "fixed before the study", and gives an insufficient panel its own sentence. Three-author citations in the handoff use "et al.". - An item rated in rounds that are not consecutive is described as such, not as rated in one round. - results gains stability_df. Other fixes: - plot() takes `type`; `which` is still accepted. - The distribution view marks a round with fewer than three experts as no decision, and bins a rating between scale points on the side the rule counts it. - The stability view scales a chi-square on its own axis and places round-pair labels at their own positions. - The kappa caveat states the interval level in use; Holey et al. (2007) are cited for what they found, and the general point to Feinstein and Cicchetti (1990); Diamond et al.'s 75% median is described accurately. - summary() names its items and its statistic and ends with the caveat. - content_report() writes a Delphi chi-square with its leading zero and df. - compare_rounds() opens with its verdict. - The Delphi vignette restores the option it changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Four reviewers checked the expert-panel commit. One regression it introduced
is fixed, and the rest tightened:
- delphi_validity() no longer stops when an item is named like a rater ID.
It builds its own judge-by-item table from long data, so the ID check is
switched off for that internal call.
- The ID check also catches the row-number column of a CSV round trip ("X",
"...1") when it counts 1, 2, 3, and more ID-like names (coder, evaluator,
pid). Its message handles several columns and says to rename a real item.
- The essentiality table prints "none" and "--" under "needed", as cvr()
does; summary() counts "Too few experts" apart from "Insufficient data".
- The essentiality figure marks a panel too small to test with a cross, and
its legend lists only what is drawn.
- A named count vector is used for item names only when the names are
complete, non-blank and unique; otherwise items are numbered as before.
- The out-of-range error says to set lo and hi, counts its columns, and
lists at most five.
- The Lawshe comparison explains the count rule when panel sizes differ,
reads its minimums from a rated item, and prints the tabled value at two
decimals whatever digits is.
- Scale points print alike on a scale with fractional bounds.
- Docs: ?cvr no longer calls the step at eight experts a defect; the ID rule
is in every affected help page; "Insufficient panel" is documented for
essentiality; Wilson et al. is cited in APA form.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on, and key definitions
Values that change:
- An item with fewer than two complete judges ("Insufficient data") no
longer enters its scale's mean HTC and HTD, where it could move the
scale's Colquitt band.
- Exactly parallel judge profiles give F = Inf and p = 0 with no
sphericity correction. Rounding noise in the error sum of squares gave
an F near 9e15 with degrees of freedom corrected by a noise epsilon.
- Item, judge, construct and target labels are compared as trimmed text,
and items keep the order of the data or of factor levels.
Attribution: Hinkin and Tracey (1999) used a one-way ANOVA with Duncan's
multiple range test. The repeated-measures ANOVA with a planned contrast
is MacKenzie, Podsakoff and Podsakoff's (2011) recommendation. The print
heading, the handoff rule and citation, and the help now say so, and the
help labels the four package choices.
Definitions: HTD is the average lead over every other construct, not the
lead over the closest one. HTC runs from 1/a to 1.
Output: the rating print states its rule and alpha; whole degrees of
freedom print whole; the anova print states the alpha and adjustment of
its contrasts; an item held back by the omnibus test alone is told so;
content_report() adds the F test and partial eta-squared; expert-judge
tables drop the level columns; the profile figure draws complete-judge
means and no gap for an item without a decision.
Handoff: schema 1 is untouched. The construct-rating p_value row carries
a note when a contrast, not the omnibus p, held the item back.
A single orbiting_r named for a construct other than the one target is
now an error in rating_validity() and sort_validity().
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…in targets Judges: - An item at the CVI criterion no longer flags its judges. Such items are reported in their own table, with the judges whose removal changes each; an item rated by three or fewer judges is "not checked". "Influential" is no longer a decision. - Fit statistics flag a judge only above Linacre's (2002) range, and only when the model scored at least fit_min_ratings (30) of their decisions. A judge below the range is described, not flagged. - The joint-maximum-likelihood correction counts judges, not items: the stretch follows J / (J - 1). Cited to Wright and Douglas (1977) and Wright (1988). - With missing ratings a judge is compared with the panel on the items they rated. - The cuts are arguments, printed and labeled as package conventions. Content structure: - Clusters come from the MDS coordinates, as in Sireci and Geisinger, so the number of dimensions matters. The help says what still differs. - A one-cell blueprint or one cluster has no adjusted Rand index. - The map's fit is printed as "raw stress", without Kruskal's labels. - A named membership must be named by the items; ari_cut is an argument. Generalizability: phi_cut is an argument; max_judges = 1 no longer stops; the interpretation states the coefficient against the criterion and says where the variance is; an insufficient design prints no empty tables. Domain: with targets, a cell far below its intended share is "Under-represented", and the thin check uses the cell's own target. Two-by-two comparators: Fisher's exact p when an expected count is below 5, and the chi-square is printed with N. qfactor_content() says its loadings are unrotated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Round order: text labels are ordered by their number only when that number is all that differs between them, so "Q4 2023" and "Q1 2024" stop with a request for a factor where ordering by the first number put the first round last. A date column is ordered by date, and one text-labeled round gets the two-rounds message. compare_rounds() compares the panel size only where the criterion depends on it (item sort, expert relevance and essentiality). A Delphi or judge comparison was told its criterion had moved when it had not. A changed setting and a changed panel are both explained. Handoff: stability rows of an item rated again after a gap carry a note naming the pair they compare; an insufficient panel gets its own rule sentence with or without a threshold; a label that says its position is not repeated. Print: the trend tables list round pairs in round order; percentages agree across the header, the interpretation and the handoff and do not depend on session options; the key no longer calls the threshold preset; kappa "can be" low, with Holey et al. cited for what they suggest; the coverage caveat says Klar et al. studied 95% intervals; the summary names items on the count lines. plot(): giving both type and which is an error; an early-ending line is labeled above its last point. A rating a hair under the cut stays on the lower side of the distribution view. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Scale means: each mean uses the items whose index rests on at least two judges, HTC by the judges who rated the intended construct (new n_target) and HTD by the complete judges. An undecided item's HTC from the whole panel is no longer dropped. - Ties: max_contrast_p is NA when a contrast has no p, and the handoff note and the printout name the tie instead of quoting a passing contrast's p. Contrasts of exactly parallel profiles give t = Inf, not rounding noise. - The two degrees of freedom of a test are formatted together, so a corrected pair prints "F(1.39, 32.00)", and Inf is never padded. - Attribution: the omnibus-then-contrast sequence is MacKenzie et al.'s; the help lists the package's own choices (Greenhouse-Geisser p, one contrast per orbiting construct, no adjustment by default, two-judge minimum, Welch contrasts between judges). The anova print names its sources on a second line; the help page is retitled. - The rating APA report fits 80 columns; the competitor and partial eta-squared stay in format = "data.frame". A missing F prints "--". - Holm adjustment is named in the rule, the note and the printout; a missing omnibus p carries a note. - Expert-judge tables are headed "means"; the summary drops level columns by judge type; the profile key takes two rows when it has five entries; the profile layout is a tested helper. - Wording: verb agreement, "every other construct", target_map error, orbiting_r name trimming. Colquitt et al. (2014), no longer cited, is removed from the references; label conversion is faster. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # R/glossary.R # inst/WORDLIST
Judges: the fragile-item note says such an item is one judge from the other side of the criterion (at it, or one short of it), which holds for every panel size. Complete ratings give raw severity bit for bit as before, and a judge exactly at a cut is not flagged. With missing ratings the bias correction counts the judges behind each item, and standard errors are scaled by the square root of the factor. The fit rule cites Linacre only for his own bounds, labels the flag and the 30 as package conventions, and names judges above the range on too few decisions. One judge prints only the verdict; Phi is not said to rest on fewer than two judges; summary() repeats the missing-ratings note; objects saved before 1.0 print again. The help states both of Linacre's bands and that applying the Wright-Douglas correction to judges is this package's choice. Structure and domain: the map's fit is printed as "distortion" (raw stress names another quantity), GOF is described as a share of the absolute eigenvalues, and average linkage is labeled the package's choice. The example similarities now place nine distinct items. A pch or col given per cell is applied by cell and kept in the key. Domain criteria state the target floor (rounded up) and the convention label; a target cell with no items is reported as not covered; similarity labels are trimmed. A cut the analyst sets is printed as such, and a value that rounds to its cut gets a third decimal. compare_rounds() ignores the settings the data decide. Large two-by-two tables no longer overflow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ions ioc() returns Rovinelli and Hambleton's index; congruence decides on it against ioc_cut (.70), with the means and margin described beside it, one row per item without targets, and handoff statistics renamed to match. The unreachable relevance tier "Support" is removed. Lynn's criterion beyond ten experts and expert_power() are labeled as package extensions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # R/expert_validity.R
…sign The scale summaries of sort_validity() and rating_validity() no longer report overall_strength, a combination of the two Colquitt et al. (2019) bands that they do not publish; the evidence sentence states each band. A caution, labeled as the package's own, is added when judges saw other than three definitions, the design behind every scale in their norms. Older IOC tests check the published index beside the mean, and the regenerated help, NEWS and vignette describe the change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ames The key defines S-CVI, states when modified kappa falls below 0, describes AC1 when it is chosen, separates severity in rating points from logits, and states the domain rules. Three-author works use et al., references carry article numbers and are cited where listed, and the vignettes build their tables with content_report(). Every plot method takes xlab, ylab, xlim, ylim and main from the caller; long item names are shortened rather than breaking a figure. The relevance report names its two intervals apart, and as.data.frame() works on every result. The walkthrough and handoff vignettes are corrected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A criterion set with ioc_cut is no longer credited to Rovinelli and Hambleton, in the print, the handoff rule or the key. A congruent item whose rival is rated as high as the target says so. Fits saved before 1.0 print with a note and are refused elsewhere, since their target_ioc held a mean. The congruence plot draws both means below the index, marks items with no index and gives the criterion's value; the verdict names items rated on their target only. The Colquitt caution is labeled as the package's own, counts the definitions each item was rated against (or that judges used), and prints under the benchmark table; the help describes the sorting and rating designs apart, and the README explains the caution its example prints. The over-ten extension is labeled wherever it is applied, with the bound of Polit and Beck's own recommendation. The index's handling of missing ratings is labeled as the package's own. The sort verdict agrees in number, and NEWS covers the changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # R/expert_validity.R # R/glossary.R
Callers can no longer override what a figure's encoding depends on (frame type, the method's axes, a map's decision symbols); a NULL keeps the method's value and unnamed arguments warn. Titles in the distribution view and the evidence profile are drawn once above the legend. Item labels are measured on the device, shortened in the middle and kept distinct, with room reserved for Delphi round labels. Report intervals are named after their estimates and a renamed column prints under its new name. The glossary keys severity by its results column, the AC1 key no longer reads 0 as chance, and the thin-coverage meaning no longer mentions targets. Three-author works read et al. throughout, help pages and vignettes cite what they list, the walkthrough's TF4 sentence and the reporting templates are corrected, and the new help topic is in the pkgdown index. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Seventh of the fix PRs from the pre-1.0 audit. It covers the keys, help pages, figures and report tables. Some returned values change, so those come first.
Values that change
tab$"95% CI"found only the first. Every interval column ofcontent_report()'s table is now named after its estimate:Psa 95% CIin a sort report,V 95% CIandI-CVI 95% CIin a relevance report, and the Delphi stability interval (kappa 95% CIby default); the level followsalpha. Code that readtab$"95% CI"from any report needs the new name. The printed table and the Markdown still head them "95% CI" beside their estimate, and a renamed column prints under its new name.as.data.frame()oncsv_binom_test(),signal_detection()andreproducibility_phi()gave two to four rows; it now gives one, with an interval asci_lowandci_high(the binomial one marked one-sided byci_sidesandci_level) and a two-by-two table as four namedn_counts.severityterm is now the logit severity; the rating-point severity it described isseverity_raw, matching theresultscolumns. New terms:agreement_ac1,S_CVI_Ave,S_CVI_UA.plot()takesxlab,ylab,xlim,ylimandmain. Most methods set these and also passed...on, so giving one stopped with "formal argument matched by multiple actual arguments". A caller's argument now replaces the method's own, except the frame type, the axes the method draws, and the decision symbols a legend keys. An unnamed argument is dropped with a warning. In the distribution views and the evidence profile,mainis drawn once above the legend andylabis ignored, because the item labels take its place; the flow diagram takes none of these.Other fixes
as.data.frame()now also works oncvi(),gtheory_content(),content_structure(),compare_rounds(),content_handoff(),sort_power(),expert_power()andpanel_agreement(), withcomponentwhere a result holds several tables; the agreement row names the interval's error rateci_alpha(?"contentvalid-data-frames").?contentvalidRno longer promises a plot for judge and domain fits, and says how to draw a domain fit's content map.content_report(); the walkthrough and reading-output vignettes are corrected; the handoff vignette's nomologR install line names CRAN as well as R-universe.The handoff
Unchanged.
How it was checked
R CMD check: 0 errors, 0 warnings, the 2 expected notes).Not in this PR
The output style shared with nomologR (headers, rounding, tables, notes) is the next PR.
🤖 Generated with Claude Code