Outcome
Create an operator-reviewed, reproducible corpus that measures whether Kai's generated facts and episodes are genuinely useful and accurate.
Requirements
- Sample authorized production exchanges and their extraction outputs without placing private corpus material in the public repository.
- Define labels for useful, incorrect, unsupported, transient, stale-on-arrival, redundant, fragmented, wrongly scoped, wrong-speaker, missed update, and episode false positive/false negative.
- Separate fact extraction, episode classification, episode generation, consolidation intent, and scope quality.
- Include changed-value, negation, preference reversal, renamed resource, completed workflow, routine acknowledgment, multi-agent, and project/global-scope cases.
- Preserve immutable corpus inputs and reviewer decisions with provenance.
- Support blind comparison of model outputs.
Acceptance
- Baseline precision, useful-memory rate, duplication rate, fragmentation rate, update-detection rate, and episode quality are reported for the current production pipeline.
- Corpus storage remains private and principal-authorized.
- The corpus is large and varied enough to prevent selecting a model from anecdotal examples.
Outcome
Create an operator-reviewed, reproducible corpus that measures whether Kai's generated facts and episodes are genuinely useful and accurate.
Requirements
Acceptance