Skip to content

Build an operator-reviewed production memory quality corpus and scoring rubric #1706

Description

@dcellison

Outcome

Create an operator-reviewed, reproducible corpus that measures whether Kai's generated facts and episodes are genuinely useful and accurate.

Requirements

  • Sample authorized production exchanges and their extraction outputs without placing private corpus material in the public repository.
  • Define labels for useful, incorrect, unsupported, transient, stale-on-arrival, redundant, fragmented, wrongly scoped, wrong-speaker, missed update, and episode false positive/false negative.
  • Separate fact extraction, episode classification, episode generation, consolidation intent, and scope quality.
  • Include changed-value, negation, preference reversal, renamed resource, completed workflow, routine acknowledgment, multi-agent, and project/global-scope cases.
  • Preserve immutable corpus inputs and reviewer decisions with provenance.
  • Support blind comparison of model outputs.

Acceptance

  • Baseline precision, useful-memory rate, duplication rate, fragmentation rate, update-detection rate, and episode quality are reported for the current production pipeline.
  • Corpus storage remains private and principal-authorized.
  • The corpus is large and varied enough to prevent selecting a model from anecdotal examples.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    architectureDesign decisions and architectural directionenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions