Skip to content

Evaluate fact and episode models against the reviewed corpus #1709

Description

@dcellison

Outcome

Select fact-extraction, episode-classification, and episode-generation models using the reviewed corpus rather than assuming a stronger or more expensive model is better.

Requirements

  • Replay identical inputs and candidate context through current cheap-tier models and approved stronger candidates.
  • Measure strict factual precision, useful recall, over-extraction, fragmentation, consolidation intent, scope, episode classification, structured-output reliability, latency, and provider failures.
  • Evaluate fact extraction and episode generation independently.
  • Produce a versioned comparison report and explicit promotion threshold.
  • Permit per-role defaults or policy-bounded overrides without changing any conversational agent's selected/default model.
  • Fail closed when the selected role model is unavailable or outside runtime authority.

Acceptance

  • The production choice is justified by recorded metrics.
  • Re-running the evaluation against the same corpus is deterministic apart from documented provider variance.
  • A model promotion has a rollback path and preserves historical model provenance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    architectureDesign decisions and architectural directionenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions