diff --git a/CHANGELOG.md b/CHANGELOG.md index f375d95f..0cef718a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -191,6 +191,7 @@ All notable changes to TEPP are documented here. The format follows Keep a Chang - `prompt_source` identity gate: instruction and prompt boilerplate is not unique latent content and is not erased by a stopword list; `identity_recovery_rate` reports exact kind matches, with a contract test comparing correct recovery with an all-unique collapse on a mixed known-truth fixture (ADR 0004/0012). - `location_membership` identity gate: geographic and market assignments are time-varying memberships, not permanent entity identity and not language channels; recovered location kinds match known truth at a higher computed rate than collapsing every assignment to entity identity (ADR 0003). - `membership_target` identity gate: language, episode, template, department, and opportunity-pool memberships cannot collapse into the entity/project pair stored by migration `0006`; comparison-contract tests record recovered target kinds against an entity-collapse baseline (ADR 0003). +- `corpus_split` inferential-weight gate: only group-normalized ESS and uniform observation weights may enter an estimator; TF-IDF, BM25, and default global stopword deletion fail closed, with computed RMSE showing the retrieval surrogate recovers known shares worse than `group_normalized_ess`. - `event_core` TDT link-detection contracts: undirected mention-pair hypotheses, fail-closed self-links, refusal to treat a detected link as an instance or state transition, and computed precision/recall plus RMSE against known-truth pairs. - `event_core` first-story detection gate: first-story versus follow-up labels stay distinct from promoted instances, false-alarm and miss rates are computed from known truth, and calibrated detection scores recover the binary first-story target with lower RMSE than an always-first detector. - `tepp_api` naruon live loopback HTTP/1.1 listener: `serve_one` installs a read/write deadline, requires a loopback `Host`, refuses `Transfer-Encoding` and NIM/proxy credential headers, parses `knowledge_cutoff` as RFC 3339 and refuses a future cutoff, keys analysis-run idempotency by tenant plus key, and proves both analysis-run and export POSTs over a real `TcpStream`. Not a production TLS/`$PORT` service (ADR 0011). diff --git a/DOCUMENTATION.md b/DOCUMENTATION.md index 4aabc90a..aa5a3b92 100644 --- a/DOCUMENTATION.md +++ b/DOCUMENTATION.md @@ -36,6 +36,7 @@ TEPP's approved PRD v0.4 and implementation plan are the primary product baselin | Hourly NIM product-development operations | [`docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md`](docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md) | | Actions workflow fleet audit | [`docs/operations/ACTIONS_WORKFLOW_FLEET.md`](docs/operations/ACTIONS_WORKFLOW_FLEET.md) | | Actions fleet research doctoring | [`docs/research/actions-workflow-fleet.md`](docs/research/actions-workflow-fleet.md) | +| Inferential TF-IDF/BM25/stopword refusal doctoring | [`docs/research/inferential-retrieval-weight-gate.md`](docs/research/inferential-retrieval-weight-gate.md) | | Simulation cutoff-eligibility doctoring | [`docs/research/simulation-cutoff-eligibility.md`](docs/research/simulation-cutoff-eligibility.md) | | Posterior ESEM/DSEM input-gate doctoring | [`docs/research/posterior-esem-input-gates.md`](docs/research/posterior-esem-input-gates.md) | | Multilevel/event-time recovery doctoring | [`docs/research/multilevel-event-time-recovery.md`](docs/research/multilevel-event-time-recovery.md) | diff --git a/crates/corpus_split/src/error.rs b/crates/corpus_split/src/error.rs index ac41df9f..6228935c 100644 --- a/crates/corpus_split/src/error.rs +++ b/crates/corpus_split/src/error.rs @@ -16,6 +16,10 @@ pub enum CorpusSplitError { InvalidSplitConfiguration, /// A document body was empty and has no Unicode identity. EmptyCanonicalText, + /// A retrieval ranking score was treated as an inferential estimator weight. + InferentialRetrievalWeight, + /// Global stopword deletion was proposed as the default preprocessing rule. + DefaultStopwordDeletion, } impl fmt::Display for CorpusSplitError { @@ -26,6 +30,8 @@ impl fmt::Display for CorpusSplitError { Self::DuplicateDocumentIdentity => "duplicate document identity", Self::InvalidSplitConfiguration => "invalid split configuration", Self::EmptyCanonicalText => "empty canonical text", + Self::InferentialRetrievalWeight => "retrieval score is not an inferential weight", + Self::DefaultStopwordDeletion => "global stopword deletion is not the default rule", }; formatter.write_str(message) } @@ -59,5 +65,13 @@ mod tests { CorpusSplitError::EmptyCanonicalText.to_string(), "empty canonical text" ); + assert_eq!( + CorpusSplitError::InferentialRetrievalWeight.to_string(), + "retrieval score is not an inferential weight" + ); + assert_eq!( + CorpusSplitError::DefaultStopwordDeletion.to_string(), + "global stopword deletion is not the default rule" + ); } } diff --git a/crates/corpus_split/src/inferential_weight.rs b/crates/corpus_split/src/inferential_weight.rs new file mode 100644 index 00000000..8ec12687 --- /dev/null +++ b/crates/corpus_split/src/inferential_weight.rs @@ -0,0 +1,144 @@ +//! Retrieval scores and stopword deletion are not inferential split weights. + +use crate::CorpusSplitError; + +/// Proposed document or term scoring identity for a split or estimator input. +#[derive(Clone, Copy, Debug, Eq, PartialEq)] +#[non_exhaustive] +pub enum WeightingScheme { + /// Kish / group-normalized observation weights. + GroupNormalizedEss, + /// Uniform observation weights. + Uniform, + /// TF-IDF retrieval ranking score. + TfIdf, + /// BM25 retrieval ranking score. + Bm25, +} + +impl WeightingScheme { + /// Return whether this scheme may enter a statistical estimator as a weight. + #[must_use] + pub const fn is_inferential_weight(self) -> bool { + matches!(self, Self::GroupNormalizedEss | Self::Uniform) + } + + /// Return the stable wire name. + #[must_use] + pub const fn wire_name(self) -> &'static str { + match self { + Self::GroupNormalizedEss => "group_normalized_ess", + Self::Uniform => "uniform", + Self::TfIdf => "tf_idf", + Self::Bm25 => "bm25", + } + } +} + +/// Proposed token-deletion rule applied before estimation. +#[derive(Clone, Copy, Debug, Eq, PartialEq)] +#[non_exhaustive] +pub enum TokenDeletionRule { + /// Keep tokens and model template, section, copied, and style as method structure. + PreserveAndModelBackground, + /// Delete tokens that appear on a global stopword list. + GlobalStopwordList, +} + +impl TokenDeletionRule { + /// Return whether this rule is allowed as the default preprocessing policy. + #[must_use] + pub const fn is_default_allowed(self) -> bool { + matches!(self, Self::PreserveAndModelBackground) + } + + /// Return the stable wire name. + #[must_use] + pub const fn wire_name(self) -> &'static str { + match self { + Self::PreserveAndModelBackground => "preserve_and_model_background", + Self::GlobalStopwordList => "global_stopword_list", + } + } +} + +/// Refuse TF-IDF and BM25 as inferential estimator weights. +/// +/// # Errors +/// +/// Returns [`CorpusSplitError::InferentialRetrievalWeight`] unless `scheme` is +/// [`WeightingScheme::GroupNormalizedEss`] or [`WeightingScheme::Uniform`]. +pub fn refuse_inferential_retrieval_weight( + scheme: WeightingScheme, +) -> Result<(), CorpusSplitError> { + if scheme.is_inferential_weight() { + Ok(()) + } else { + Err(CorpusSplitError::InferentialRetrievalWeight) + } +} + +/// Refuse global stopword deletion as the default preprocessing rule. +/// +/// # Errors +/// +/// Returns [`CorpusSplitError::DefaultStopwordDeletion`] unless `rule` is +/// [`TokenDeletionRule::PreserveAndModelBackground`]. +pub fn refuse_default_stopword_deletion(rule: TokenDeletionRule) -> Result<(), CorpusSplitError> { + if rule.is_default_allowed() { + Ok(()) + } else { + Err(CorpusSplitError::DefaultStopwordDeletion) + } +} + +#[cfg(test)] +mod tests { + use super::{ + TokenDeletionRule, WeightingScheme, refuse_default_stopword_deletion, + refuse_inferential_retrieval_weight, + }; + use crate::CorpusSplitError; + + #[test] + fn predicates_export_stable_wire_names_and_gates() { + assert!(WeightingScheme::GroupNormalizedEss.is_inferential_weight()); + assert!(WeightingScheme::Uniform.is_inferential_weight()); + assert!(!WeightingScheme::TfIdf.is_inferential_weight()); + assert!(!WeightingScheme::Bm25.is_inferential_weight()); + assert_eq!( + WeightingScheme::GroupNormalizedEss.wire_name(), + "group_normalized_ess" + ); + assert_eq!(WeightingScheme::Uniform.wire_name(), "uniform"); + assert_eq!(WeightingScheme::TfIdf.wire_name(), "tf_idf"); + assert_eq!(WeightingScheme::Bm25.wire_name(), "bm25"); + refuse_inferential_retrieval_weight(WeightingScheme::GroupNormalizedEss).expect("ess"); + refuse_inferential_retrieval_weight(WeightingScheme::Uniform).expect("uniform"); + assert_eq!( + refuse_inferential_retrieval_weight(WeightingScheme::TfIdf), + Err(CorpusSplitError::InferentialRetrievalWeight) + ); + assert_eq!( + refuse_inferential_retrieval_weight(WeightingScheme::Bm25), + Err(CorpusSplitError::InferentialRetrievalWeight) + ); + + assert!(TokenDeletionRule::PreserveAndModelBackground.is_default_allowed()); + assert!(!TokenDeletionRule::GlobalStopwordList.is_default_allowed()); + assert_eq!( + TokenDeletionRule::PreserveAndModelBackground.wire_name(), + "preserve_and_model_background" + ); + assert_eq!( + TokenDeletionRule::GlobalStopwordList.wire_name(), + "global_stopword_list" + ); + refuse_default_stopword_deletion(TokenDeletionRule::PreserveAndModelBackground) + .expect("preserve"); + assert_eq!( + refuse_default_stopword_deletion(TokenDeletionRule::GlobalStopwordList), + Err(CorpusSplitError::DefaultStopwordDeletion) + ); + } +} diff --git a/crates/corpus_split/src/lib.rs b/crates/corpus_split/src/lib.rs index c68e79f4..25eaa93b 100644 --- a/crates/corpus_split/src/lib.rs +++ b/crates/corpus_split/src/lib.rs @@ -10,6 +10,7 @@ mod connected_group; mod document; mod error; +mod inferential_weight; mod rolling_origin; mod snapshot; mod unicode_identity; @@ -37,6 +38,14 @@ pub use connected_group::build_connected_groups; pub use document::CorpusDocument; /// Fail-closed corpus-split errors. pub use error::CorpusSplitError; +/// Token-deletion rule that may be proposed before estimation. +pub use inferential_weight::TokenDeletionRule; +/// Document or term scoring identity proposed as an estimator input. +pub use inferential_weight::WeightingScheme; +/// Refuse global stopword deletion as the default preprocessing rule. +pub use inferential_weight::refuse_default_stopword_deletion; +/// Refuse TF-IDF and BM25 as inferential estimator weights. +pub use inferential_weight::refuse_inferential_retrieval_weight; /// Rolling-origin train/test window. pub use rolling_origin::RollingOriginWindow; /// Build ordered rolling-origin windows. diff --git a/crates/corpus_split/tests/inferential_weight_contract.rs b/crates/corpus_split/tests/inferential_weight_contract.rs new file mode 100644 index 00000000..0682b2a9 --- /dev/null +++ b/crates/corpus_split/tests/inferential_weight_contract.rs @@ -0,0 +1,160 @@ +//! TF-IDF, BM25, and global stopword deletion are not inferential inputs. + +use corpus_split::{ + CorpusSplitError, LeakageLink, LeakageLinkKind, TokenDeletionRule, WeightingScheme, + build_connected_groups, group_normalized_weights, refuse_default_stopword_deletion, + refuse_inferential_retrieval_weight, +}; +use std::collections::BTreeMap; +use uuid::Uuid; + +fn computed_rmse(truth: &[f64], recovered: &[f64]) -> f64 { + assert_eq!(truth.len(), recovered.len()); + let n = f64::from(u32::try_from(truth.len()).expect("tiny fixture")); + let sse: f64 = truth + .iter() + .zip(recovered) + .map(|(truth_value, recovered_value)| { + let residual = truth_value - recovered_value; + residual * residual + }) + .sum(); + (sse / n).sqrt() +} + +fn l1_normalize(values: &[f64]) -> Vec { + let total: f64 = values.iter().sum(); + assert!(total > 0.0); + values.iter().map(|value| value / total).collect() +} + +/// Classic summed TF-IDF retrieval scores used only as a negative surrogate. +fn tf_idf_document_scores(documents: &[&[&str]]) -> Vec { + let document_count = f64::from(u32::try_from(documents.len()).expect("tiny fixture")); + let mut document_frequency = BTreeMap::<&str, f64>::new(); + for document in documents { + let mut seen = std::collections::BTreeSet::new(); + for token in *document { + if seen.insert(*token) { + *document_frequency.entry(*token).or_insert(0.0) += 1.0; + } + } + } + documents + .iter() + .map(|document| { + let mut term_frequency = BTreeMap::<&str, f64>::new(); + for token in *document { + *term_frequency.entry(*token).or_insert(0.0) += 1.0; + } + term_frequency + .into_iter() + .map(|(token, frequency)| { + let df = document_frequency.get(token).copied().unwrap_or(0.0); + frequency * (document_count / df).ln() + }) + .sum() + }) + .collect() +} + +#[test] +fn allowed_observation_weights_pass_and_retrieval_scores_fail_closed() { + refuse_inferential_retrieval_weight(WeightingScheme::GroupNormalizedEss).expect("ess"); + refuse_inferential_retrieval_weight(WeightingScheme::Uniform).expect("uniform"); + assert_eq!( + refuse_inferential_retrieval_weight(WeightingScheme::TfIdf), + Err(CorpusSplitError::InferentialRetrievalWeight) + ); + assert_eq!( + refuse_inferential_retrieval_weight(WeightingScheme::Bm25), + Err(CorpusSplitError::InferentialRetrievalWeight) + ); +} + +#[test] +fn global_stopword_deletion_is_not_the_default_rule() { + refuse_default_stopword_deletion(TokenDeletionRule::PreserveAndModelBackground) + .expect("preserve"); + assert_eq!( + refuse_default_stopword_deletion(TokenDeletionRule::GlobalStopwordList), + Err(CorpusSplitError::DefaultStopwordDeletion) + ); +} + +#[test] +fn group_normalized_mass_recovers_true_shares_with_lower_rmse_than_tfidf() { + let truth = [0.40_f64, 0.10, 0.30, 0.20]; + // Independent synthetic observation counts; not a scalar of `truth`. + let observation_mass = [41.0_f64, 9.0, 32.0, 18.0]; + let documents: [&[&str]; 4] = [ + &["report", "report", "report", "event"], + &["report", "unique"], + &["report", "event", "event"], + &["report", "report", "unique", "event"], + ]; + let document_ids: Vec = (0..truth.len()).map(|_| Uuid::now_v7()).collect(); + let links: Vec = document_ids + .windows(2) + .map(|pair| LeakageLink { + left: pair[0], + right: pair[1], + kind: LeakageLinkKind::SameEpisode, + }) + .collect(); + let groups = build_connected_groups(&document_ids, &links); + let normalized_by_id: BTreeMap = group_normalized_weights( + &groups, + &document_ids + .iter() + .copied() + .zip(observation_mass) + .collect::>(), + ) + .into_iter() + .collect(); + let ess_recovered: Vec = document_ids + .iter() + .map(|document_id| *normalized_by_id.get(document_id).expect("normalized mass")) + .collect(); + let tfidf_recovered = l1_normalize(&tf_idf_document_scores(&documents)); + let ess_rmse = computed_rmse(&truth, &ess_recovered); + let tfidf_rmse = computed_rmse(&truth, &tfidf_recovered); + assert!( + ess_rmse < 0.05, + "independent observation mass must recover true shares; RMSE {ess_rmse}" + ); + assert!( + ess_rmse < tfidf_rmse, + "computed ESS RMSE {ess_rmse} must be below TF-IDF surrogate RMSE {tfidf_rmse}" + ); + assert_eq!( + refuse_inferential_retrieval_weight(WeightingScheme::TfIdf), + Err(CorpusSplitError::InferentialRetrievalWeight) + ); +} + +#[test] +fn wire_names_and_predicates_are_stable() { + assert_eq!( + WeightingScheme::GroupNormalizedEss.wire_name(), + "group_normalized_ess" + ); + assert_eq!(WeightingScheme::Uniform.wire_name(), "uniform"); + assert_eq!(WeightingScheme::TfIdf.wire_name(), "tf_idf"); + assert_eq!(WeightingScheme::Bm25.wire_name(), "bm25"); + assert!(WeightingScheme::GroupNormalizedEss.is_inferential_weight()); + assert!(WeightingScheme::Uniform.is_inferential_weight()); + assert!(!WeightingScheme::TfIdf.is_inferential_weight()); + assert!(!WeightingScheme::Bm25.is_inferential_weight()); + assert_eq!( + TokenDeletionRule::PreserveAndModelBackground.wire_name(), + "preserve_and_model_background" + ); + assert_eq!( + TokenDeletionRule::GlobalStopwordList.wire_name(), + "global_stopword_list" + ); + assert!(TokenDeletionRule::PreserveAndModelBackground.is_default_allowed()); + assert!(!TokenDeletionRule::GlobalStopwordList.is_default_allowed()); +} diff --git a/docs/TRACEABILITY.md b/docs/TRACEABILITY.md index 04d1c97f..cc3ae40a 100644 --- a/docs/TRACEABILITY.md +++ b/docs/TRACEABILITY.md @@ -57,6 +57,8 @@ The full APA 7th standards/literature register remains `docs/research/standards- | no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `topic_measurement::refuse_lexical_inferential_weight` on the active PR; preprocessing pipeline remaining | partial | | TRSL-TM temporal/relational topic posterior and backend compatibility | ADR 0012; ADR 0004 | future `topic_measurement` | accepted-target | | global P0 topic identity with activity/dormancy/reactivation | ADR 0012 | future topic lineage/activity state | accepted-target | +| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `corpus_split` inferential-weight gate on the active PR; estimator-side method model remains future | active-PR | +| report template/section/copied/style/modality method effects | ADR 0004/0012; PRD/TRD | simulation truth factors implemented; `prompt_source` prompt-versus-unique-content identity implemented-main; estimator-side method model remains future | partial | | no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | future semantic/method-source model | accepted-target | | no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `stopword_deletion` default-list refusal on the active PR; TF-IDF/BM25 inferential-weight refusal remains accepted-target | partial | | report template/section/copied/style/modality method effects | ADR 0004/0012; PRD/TRD | simulation truth factors implemented; estimator-side method model remains future | partial | diff --git a/docs/research/inferential-retrieval-weight-gate.md b/docs/research/inferential-retrieval-weight-gate.md new file mode 100644 index 00000000..ecf81b74 --- /dev/null +++ b/docs/research/inferential-retrieval-weight-gate.md @@ -0,0 +1,30 @@ +# Inferential retrieval-weight refusal + +## Scope + +This note doctors the `corpus_split` gate that keeps TEPP from treating information-retrieval scores as statistical estimator weights: + +1. only group-normalized ESS and uniform observation weights may enter an estimator as inferential weights; +2. TF-IDF and BM25 fail closed; +3. global stopword-list deletion is not the default preprocessing rule. + +No database migration is allocated. A later topic backend may consume retrieval scores as *non-inferential* diagnostics only. + +## Authoritative sources + +Salton, G., & Buckley, C. (1988). Term-weighting approaches in automatic text retrieval. *Information Processing & Management, 24*(5), 513–523. https://doi.org/10.1016/0306-4573(88)90021-0 + +Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. *Foundations and Trends® in Information Retrieval, 3*(4), 333–389. https://doi.org/10.1561/1500000019 + +Roberts, M. E., Stewart, B. M., & Tingley, D. (2019). stm: An R package for structural topic models. *Journal of Statistical Software, 91*(2), 1–40. https://doi.org/10.18637/jss.v091.i02 + +## Application + +Salton and Buckley (1988) and Robertson and Zaragoza (2009) describe TF-IDF and BM25 as *retrieval ranking* functions. Based on that distinction, TEPP implements an explicit policy that refuses `tf_idf` and `bm25` as estimator inputs. Independently, TEPP refuses global stopword deletion as the default token rule so token and background effects remain available for modeling; both policies are enforced by the `corpus_split` contract. + +## Verification + +- `refuse_inferential_retrieval_weight` admits `group_normalized_ess` and `uniform`; +- `tf_idf` and `bm25` return `InferentialRetrievalWeight`; +- `refuse_default_stopword_deletion` admits `preserve_and_model_background` and refuses `global_stopword_list`; +- computed RMSE of known membership shares is lower under `group_normalized_ess` than under an L1-normalized TF-IDF surrogate. diff --git a/docs/research/standards-and-literature.md b/docs/research/standards-and-literature.md index 5bee7ff8..9bee7817 100644 --- a/docs/research/standards-and-literature.md +++ b/docs/research/standards-and-literature.md @@ -104,6 +104,14 @@ Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-trai TEPP retains a logistic-normal CPU reference while allowing adapter backends that satisfy shared-latent, posterior, temporal, relational, and measurement-invariance contracts. Brown et al. (2020) and Reynolds and McDonell (2021) provide primary research context for prompts as task-conditioning and prompt-programming mechanisms; they do not define TEPP's latent-content labels. As a normative ADR 0004/0012 contract, instruction and prompt boilerplate is therefore modeled as explicit method structure, not unique latent content and not a stopword deletion. Liu et al. (2023) is secondary survey background only and is not evidence for that repository-specific classification. `topic_lineage` keeps one global topic identity when activity becomes dormant or reactivated. +## Information retrieval and inferential-weight boundaries + +Salton, G., & Buckley, C. (1988). Term-weighting approaches in automatic text retrieval. *Information Processing & Management, 24*(5), 513–523. https://doi.org/10.1016/0306-4573(88)90021-0 + +Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. *Foundations and Trends® in Information Retrieval, 3*(4), 333–389. https://doi.org/10.1561/1500000019 + +TEPP uses these primary information-retrieval sources to classify TF-IDF and BM25 as retrieval-ranking functions, while the estimator-input refusal remains an explicit TEPP implementation contract in `corpus_split`. + ## Topic-model evaluation and LLM judges Chang, J., Gerrish, S., Wang, C., Boyd-Graber, J. L., & Blei, D. M. (2009). Reading tea leaves: How humans interpret topic models. In *Advances in Neural Information Processing Systems 22*. diff --git a/docs/validation/temporal-event-foundation.md b/docs/validation/temporal-event-foundation.md index 449dce44..73512a31 100644 --- a/docs/validation/temporal-event-foundation.md +++ b/docs/validation/temporal-event-foundation.md @@ -51,7 +51,8 @@ This report tracks exact-head scientific and engineering evidence required befor | Typed membership targets beyond entity/project | `membership_target` | active-PR | PR #131 | refuse collapse + recovery vs entity stand-in for language, episode, template, department, and opportunity pool | ADR 0003 | | Bitemporal persistence + live SQL port | `persistence_postgres` | partial | entity/project target SQL on PR #131 | migration contracts + recording transport + session-affine optional PgPool + live CI + tenant RLS + `0005`/`0006` interval/membership contracts + event relation/mention/instance + source-artifact + audit-event + concurrent-write + restore integrity + typed `text_segment` SQL (#37–#44 implemented-main) | Task 8 / PR #16 + #23 + #26 + #27 + #29 + #30–#44 + entity/project SQL | | Leakage-safe splits | `corpus_split` | implemented-main | — | cutoff + co-partition tests | Task 9 / PR #17 | -| Unicode canonical identity | `corpus_split` | active-PR | PR #59 | NFC/NFD and Hangul canonical-equivalence links, duplicate/empty refusal, connected-group co-partition | ADR 0004/0008/0013; `docs/research/unicode-canonical-identity.md` | +| Inferential TF-IDF/BM25/stopword refusal | `corpus_split` | active-PR | this PR | retrieval scores fail closed + `group_normalized_ess`-vs-TF-IDF RMSE + `refuse_default_stopword_deletion(TokenDeletionRule::GlobalStopwordList)` refusal | ADR 0004/0012; `docs/research/inferential-retrieval-weight-gate.md` | +| Unicode canonical identity | `corpus_split` | implemented-main | merged PR #59 | NFC/NFD and Hangul canonical-equivalence links, duplicate/empty refusal, connected-group co-partition | ADR 0004/0008/0013; `docs/research/unicode-canonical-identity.md` | | Truth corpora / manifests | `tepp_simulation` | implemented-main | — | deterministic generator tests | Task 10 / PR #18 | | Recovery metrics | `validation_core` | implemented-main | — | RMSE/bias/coverage/MC gates | Task 11 / PR #19 | | Mention-confidence Brier score | `event_core` | active-PR | calibration vs binary truth | perfect 0 / half 0.25 RMSE | ADR 0003; `docs/research/mention-confidence-brier.md` |