Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -191,6 +191,7 @@ All notable changes to TEPP are documented here. The format follows Keep a Chang
- `prompt_source` identity gate: instruction and prompt boilerplate is not unique latent content and is not erased by a stopword list; `identity_recovery_rate` reports exact kind matches, with a contract test comparing correct recovery with an all-unique collapse on a mixed known-truth fixture (ADR 0004/0012).
- `location_membership` identity gate: geographic and market assignments are time-varying memberships, not permanent entity identity and not language channels; recovered location kinds match known truth at a higher computed rate than collapsing every assignment to entity identity (ADR 0003).
- `membership_target` identity gate: language, episode, template, department, and opportunity-pool memberships cannot collapse into the entity/project pair stored by migration `0006`; comparison-contract tests record recovered target kinds against an entity-collapse baseline (ADR 0003).
- `corpus_split` inferential-weight gate: only group-normalized ESS and uniform observation weights may enter an estimator; TF-IDF, BM25, and default global stopword deletion fail closed, with computed RMSE showing the retrieval surrogate recovers known shares worse than `group_normalized_ess`.
- `event_core` TDT link-detection contracts: undirected mention-pair hypotheses, fail-closed self-links, refusal to treat a detected link as an instance or state transition, and computed precision/recall plus RMSE against known-truth pairs.
- `event_core` first-story detection gate: first-story versus follow-up labels stay distinct from promoted instances, false-alarm and miss rates are computed from known truth, and calibrated detection scores recover the binary first-story target with lower RMSE than an always-first detector.
- `tepp_api` naruon live loopback HTTP/1.1 listener: `serve_one` installs a read/write deadline, requires a loopback `Host`, refuses `Transfer-Encoding` and NIM/proxy credential headers, parses `knowledge_cutoff` as RFC 3339 and refuses a future cutoff, keys analysis-run idempotency by tenant plus key, and proves both analysis-run and export POSTs over a real `TcpStream`. Not a production TLS/`$PORT` service (ADR 0011).
Expand Down
1 change: 1 addition & 0 deletions DOCUMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ TEPP's approved PRD v0.4 and implementation plan are the primary product baselin
| Hourly NIM product-development operations | [`docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md`](docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md) |
| Actions workflow fleet audit | [`docs/operations/ACTIONS_WORKFLOW_FLEET.md`](docs/operations/ACTIONS_WORKFLOW_FLEET.md) |
| Actions fleet research doctoring | [`docs/research/actions-workflow-fleet.md`](docs/research/actions-workflow-fleet.md) |
| Inferential TF-IDF/BM25/stopword refusal doctoring | [`docs/research/inferential-retrieval-weight-gate.md`](docs/research/inferential-retrieval-weight-gate.md) |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📝 Info: Duplicated documentation map diverges after one-sided edit

The file contains the whole documentation map twice. The new doctoring row is added only to the first copy; the second copy near line 80 is left without it, so the two tables now diverge. The duplication itself is pre-existing.

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

| Simulation cutoff-eligibility doctoring | [`docs/research/simulation-cutoff-eligibility.md`](docs/research/simulation-cutoff-eligibility.md) |
| Posterior ESEM/DSEM input-gate doctoring | [`docs/research/posterior-esem-input-gates.md`](docs/research/posterior-esem-input-gates.md) |
| Multilevel/event-time recovery doctoring | [`docs/research/multilevel-event-time-recovery.md`](docs/research/multilevel-event-time-recovery.md) |
Expand Down
14 changes: 14 additions & 0 deletions crates/corpus_split/src/error.rs
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,10 @@ pub enum CorpusSplitError {
InvalidSplitConfiguration,
/// A document body was empty and has no Unicode identity.
EmptyCanonicalText,
/// A retrieval ranking score was treated as an inferential estimator weight.
InferentialRetrievalWeight,
/// Global stopword deletion was proposed as the default preprocessing rule.
DefaultStopwordDeletion,
}

impl fmt::Display for CorpusSplitError {
Expand All @@ -26,6 +30,8 @@ impl fmt::Display for CorpusSplitError {
Self::DuplicateDocumentIdentity => "duplicate document identity",
Self::InvalidSplitConfiguration => "invalid split configuration",
Self::EmptyCanonicalText => "empty canonical text",
Self::InferentialRetrievalWeight => "retrieval score is not an inferential weight",
Self::DefaultStopwordDeletion => "global stopword deletion is not the default rule",
};
formatter.write_str(message)
}
Expand Down Expand Up @@ -59,5 +65,13 @@ mod tests {
CorpusSplitError::EmptyCanonicalText.to_string(),
"empty canonical text"
);
assert_eq!(
CorpusSplitError::InferentialRetrievalWeight.to_string(),
"retrieval score is not an inferential weight"
);
assert_eq!(
CorpusSplitError::DefaultStopwordDeletion.to_string(),
"global stopword deletion is not the default rule"
);
}
}
144 changes: 144 additions & 0 deletions crates/corpus_split/src/inferential_weight.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,144 @@
//! Retrieval scores and stopword deletion are not inferential split weights.

use crate::CorpusSplitError;

/// Proposed document or term scoring identity for a split or estimator input.
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
#[non_exhaustive]
pub enum WeightingScheme {
/// Kish / group-normalized observation weights.
GroupNormalizedEss,
/// Uniform observation weights.
Uniform,
/// TF-IDF retrieval ranking score.
TfIdf,
/// BM25 retrieval ranking score.
Bm25,
}

impl WeightingScheme {
/// Return whether this scheme may enter a statistical estimator as a weight.
#[must_use]
pub const fn is_inferential_weight(self) -> bool {
matches!(self, Self::GroupNormalizedEss | Self::Uniform)
}

/// Return the stable wire name.
#[must_use]
pub const fn wire_name(self) -> &'static str {
match self {
Self::GroupNormalizedEss => "group_normalized_ess",
Self::Uniform => "uniform",
Self::TfIdf => "tf_idf",
Self::Bm25 => "bm25",
}
}
}

/// Proposed token-deletion rule applied before estimation.
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
#[non_exhaustive]
pub enum TokenDeletionRule {
/// Keep tokens and model template, section, copied, and style as method structure.
PreserveAndModelBackground,
/// Delete tokens that appear on a global stopword list.
GlobalStopwordList,
}

impl TokenDeletionRule {
/// Return whether this rule is allowed as the default preprocessing policy.
#[must_use]
pub const fn is_default_allowed(self) -> bool {
matches!(self, Self::PreserveAndModelBackground)
}

/// Return the stable wire name.
#[must_use]
pub const fn wire_name(self) -> &'static str {
match self {
Self::PreserveAndModelBackground => "preserve_and_model_background",
Self::GlobalStopwordList => "global_stopword_list",
}
}
}

/// Refuse TF-IDF and BM25 as inferential estimator weights.
///
/// # Errors
///
/// Returns [`CorpusSplitError::InferentialRetrievalWeight`] unless `scheme` is
/// [`WeightingScheme::GroupNormalizedEss`] or [`WeightingScheme::Uniform`].
pub fn refuse_inferential_retrieval_weight(
scheme: WeightingScheme,
) -> Result<(), CorpusSplitError> {
if scheme.is_inferential_weight() {
Ok(())
} else {
Err(CorpusSplitError::InferentialRetrievalWeight)
}
}

/// Refuse global stopword deletion as the default preprocessing rule.
///
/// # Errors
///
/// Returns [`CorpusSplitError::DefaultStopwordDeletion`] unless `rule` is
/// [`TokenDeletionRule::PreserveAndModelBackground`].
pub fn refuse_default_stopword_deletion(rule: TokenDeletionRule) -> Result<(), CorpusSplitError> {
if rule.is_default_allowed() {
Ok(())
} else {
Err(CorpusSplitError::DefaultStopwordDeletion)
}
}

#[cfg(test)]
mod tests {
use super::{
TokenDeletionRule, WeightingScheme, refuse_default_stopword_deletion,
refuse_inferential_retrieval_weight,
};
use crate::CorpusSplitError;

#[test]
fn predicates_export_stable_wire_names_and_gates() {
assert!(WeightingScheme::GroupNormalizedEss.is_inferential_weight());
assert!(WeightingScheme::Uniform.is_inferential_weight());
assert!(!WeightingScheme::TfIdf.is_inferential_weight());
assert!(!WeightingScheme::Bm25.is_inferential_weight());
assert_eq!(
WeightingScheme::GroupNormalizedEss.wire_name(),
"group_normalized_ess"
);
assert_eq!(WeightingScheme::Uniform.wire_name(), "uniform");
assert_eq!(WeightingScheme::TfIdf.wire_name(), "tf_idf");
assert_eq!(WeightingScheme::Bm25.wire_name(), "bm25");
refuse_inferential_retrieval_weight(WeightingScheme::GroupNormalizedEss).expect("ess");
refuse_inferential_retrieval_weight(WeightingScheme::Uniform).expect("uniform");
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::TfIdf),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::Bm25),
Err(CorpusSplitError::InferentialRetrievalWeight)
);

assert!(TokenDeletionRule::PreserveAndModelBackground.is_default_allowed());
assert!(!TokenDeletionRule::GlobalStopwordList.is_default_allowed());
assert_eq!(
TokenDeletionRule::PreserveAndModelBackground.wire_name(),
"preserve_and_model_background"
);
assert_eq!(
TokenDeletionRule::GlobalStopwordList.wire_name(),
"global_stopword_list"
);
refuse_default_stopword_deletion(TokenDeletionRule::PreserveAndModelBackground)
.expect("preserve");
assert_eq!(
refuse_default_stopword_deletion(TokenDeletionRule::GlobalStopwordList),
Err(CorpusSplitError::DefaultStopwordDeletion)
);
}
}
9 changes: 9 additions & 0 deletions crates/corpus_split/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
mod connected_group;
mod document;
mod error;
mod inferential_weight;
mod rolling_origin;
mod snapshot;
mod unicode_identity;
Expand Down Expand Up @@ -37,6 +38,14 @@ pub use connected_group::build_connected_groups;
pub use document::CorpusDocument;
/// Fail-closed corpus-split errors.
pub use error::CorpusSplitError;
/// Token-deletion rule that may be proposed before estimation.
pub use inferential_weight::TokenDeletionRule;
/// Document or term scoring identity proposed as an estimator input.
pub use inferential_weight::WeightingScheme;
/// Refuse global stopword deletion as the default preprocessing rule.
pub use inferential_weight::refuse_default_stopword_deletion;
/// Refuse TF-IDF and BM25 as inferential estimator weights.
pub use inferential_weight::refuse_inferential_retrieval_weight;
/// Rolling-origin train/test window.
pub use rolling_origin::RollingOriginWindow;
/// Build ordered rolling-origin windows.
Expand Down
160 changes: 160 additions & 0 deletions crates/corpus_split/tests/inferential_weight_contract.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
//! TF-IDF, BM25, and global stopword deletion are not inferential inputs.

use corpus_split::{
CorpusSplitError, LeakageLink, LeakageLinkKind, TokenDeletionRule, WeightingScheme,
build_connected_groups, group_normalized_weights, refuse_default_stopword_deletion,
refuse_inferential_retrieval_weight,
};
use std::collections::BTreeMap;
use uuid::Uuid;

fn computed_rmse(truth: &[f64], recovered: &[f64]) -> f64 {
assert_eq!(truth.len(), recovered.len());
let n = f64::from(u32::try_from(truth.len()).expect("tiny fixture"));
let sse: f64 = truth
.iter()
.zip(recovered)
.map(|(truth_value, recovered_value)| {
let residual = truth_value - recovered_value;
residual * residual
})
.sum();
(sse / n).sqrt()
}

fn l1_normalize(values: &[f64]) -> Vec<f64> {
let total: f64 = values.iter().sum();
assert!(total > 0.0);
values.iter().map(|value| value / total).collect()
}

/// Classic summed TF-IDF retrieval scores used only as a negative surrogate.
fn tf_idf_document_scores(documents: &[&[&str]]) -> Vec<f64> {
let document_count = f64::from(u32::try_from(documents.len()).expect("tiny fixture"));
let mut document_frequency = BTreeMap::<&str, f64>::new();
for document in documents {
let mut seen = std::collections::BTreeSet::new();
for token in *document {
if seen.insert(*token) {
*document_frequency.entry(*token).or_insert(0.0) += 1.0;
}
}
}
documents
.iter()
.map(|document| {
let mut term_frequency = BTreeMap::<&str, f64>::new();
for token in *document {
*term_frequency.entry(*token).or_insert(0.0) += 1.0;
}
term_frequency
.into_iter()
.map(|(token, frequency)| {
let df = document_frequency.get(token).copied().unwrap_or(0.0);
frequency * (document_count / df).ln()
})
.sum()
})
.collect()
}

#[test]
fn allowed_observation_weights_pass_and_retrieval_scores_fail_closed() {
refuse_inferential_retrieval_weight(WeightingScheme::GroupNormalizedEss).expect("ess");
refuse_inferential_retrieval_weight(WeightingScheme::Uniform).expect("uniform");
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::TfIdf),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::Bm25),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
}

#[test]
fn global_stopword_deletion_is_not_the_default_rule() {
refuse_default_stopword_deletion(TokenDeletionRule::PreserveAndModelBackground)
.expect("preserve");
assert_eq!(
refuse_default_stopword_deletion(TokenDeletionRule::GlobalStopwordList),
Err(CorpusSplitError::DefaultStopwordDeletion)
);
}

#[test]
fn group_normalized_mass_recovers_true_shares_with_lower_rmse_than_tfidf() {
let truth = [0.40_f64, 0.10, 0.30, 0.20];
// Independent synthetic observation counts; not a scalar of `truth`.
let observation_mass = [41.0_f64, 9.0, 32.0, 18.0];
let documents: [&[&str]; 4] = [
&["report", "report", "report", "event"],
&["report", "unique"],
&["report", "event", "event"],
&["report", "report", "unique", "event"],
];
let document_ids: Vec<Uuid> = (0..truth.len()).map(|_| Uuid::now_v7()).collect();
let links: Vec<LeakageLink> = document_ids
.windows(2)
.map(|pair| LeakageLink {
left: pair[0],
right: pair[1],
kind: LeakageLinkKind::SameEpisode,
})
.collect();
let groups = build_connected_groups(&document_ids, &links);
let normalized_by_id: BTreeMap<Uuid, f64> = group_normalized_weights(
&groups,
&document_ids
.iter()
.copied()
.zip(observation_mass)
.collect::<Vec<_>>(),
)
.into_iter()
.collect();
let ess_recovered: Vec<f64> = document_ids
.iter()
.map(|document_id| *normalized_by_id.get(document_id).expect("normalized mass"))
.collect();
let tfidf_recovered = l1_normalize(&tf_idf_document_scores(&documents));
let ess_rmse = computed_rmse(&truth, &ess_recovered);
let tfidf_rmse = computed_rmse(&truth, &tfidf_recovered);
assert!(
ess_rmse < 0.05,
"independent observation mass must recover true shares; RMSE {ess_rmse}"
Comment thread
seonghobae marked this conversation as resolved.
);
assert!(
ess_rmse < tfidf_rmse,
"computed ESS RMSE {ess_rmse} must be below TF-IDF surrogate RMSE {tfidf_rmse}"
);
Comment thread
coderabbitai[bot] marked this conversation as resolved.
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::TfIdf),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
}

#[test]
fn wire_names_and_predicates_are_stable() {
assert_eq!(
WeightingScheme::GroupNormalizedEss.wire_name(),
"group_normalized_ess"
);
assert_eq!(WeightingScheme::Uniform.wire_name(), "uniform");
assert_eq!(WeightingScheme::TfIdf.wire_name(), "tf_idf");
assert_eq!(WeightingScheme::Bm25.wire_name(), "bm25");
assert!(WeightingScheme::GroupNormalizedEss.is_inferential_weight());
assert!(WeightingScheme::Uniform.is_inferential_weight());
assert!(!WeightingScheme::TfIdf.is_inferential_weight());
assert!(!WeightingScheme::Bm25.is_inferential_weight());
assert_eq!(
TokenDeletionRule::PreserveAndModelBackground.wire_name(),
"preserve_and_model_background"
);
assert_eq!(
TokenDeletionRule::GlobalStopwordList.wire_name(),
"global_stopword_list"
);
assert!(TokenDeletionRule::PreserveAndModelBackground.is_default_allowed());
assert!(!TokenDeletionRule::GlobalStopwordList.is_default_allowed());
}
2 changes: 2 additions & 0 deletions docs/TRACEABILITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,8 @@ The full APA 7th standards/literature register remains `docs/research/standards-
| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `topic_measurement::refuse_lexical_inferential_weight` on the active PR; preprocessing pipeline remaining | partial |
| TRSL-TM temporal/relational topic posterior and backend compatibility | ADR 0012; ADR 0004 | future `topic_measurement` | accepted-target |
| global P0 topic identity with activity/dormancy/reactivation | ADR 0012 | future topic lineage/activity state | accepted-target |
| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `corpus_split` inferential-weight gate on the active PR; estimator-side method model remains future | active-PR |
| report template/section/copied/style/modality method effects | ADR 0004/0012; PRD/TRD | simulation truth factors implemented; `prompt_source` prompt-versus-unique-content identity implemented-main; estimator-side method model remains future | partial |
| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | future semantic/method-source model | accepted-target |
| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `stopword_deletion` default-list refusal on the active PR; TF-IDF/BM25 inferential-weight refusal remains accepted-target | partial |
Comment on lines +60 to 63

@devin-ai-integration devin-ai-integration Bot Aug 25, 2026

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📝 Info: Stale traceability rows contradict new gate status

The added row marks the corpus_split inferential-weight gate active-PR, but nearby existing rows still say the TF-IDF/BM25 refusal is remaining/accepted-target. These older rows are now stale and contradictory.

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

| report template/section/copied/style/modality method effects | ADR 0004/0012; PRD/TRD | simulation truth factors implemented; estimator-side method model remains future | partial |
Expand Down
Loading
Loading