Outcome
Select fact-extraction, episode-classification, and episode-generation models using the reviewed corpus rather than assuming a stronger or more expensive model is better.
Requirements
- Replay identical inputs and candidate context through current cheap-tier models and approved stronger candidates.
- Measure strict factual precision, useful recall, over-extraction, fragmentation, consolidation intent, scope, episode classification, structured-output reliability, latency, and provider failures.
- Evaluate fact extraction and episode generation independently.
- Produce a versioned comparison report and explicit promotion threshold.
- Permit per-role defaults or policy-bounded overrides without changing any conversational agent's selected/default model.
- Fail closed when the selected role model is unavailable or outside runtime authority.
Acceptance
- The production choice is justified by recorded metrics.
- Re-running the evaluation against the same corpus is deterministic apart from documented provider variance.
- A model promotion has a rollback path and preserves historical model provenance.
Outcome
Select fact-extraction, episode-classification, and episode-generation models using the reviewed corpus rather than assuming a stronger or more expensive model is better.
Requirements
Acceptance