馃И test: Evaluate Context Selection With Offline Replay - #332
Open
lia-by-librechat[bot] wants to merge 1 commit into
Open
lia-by-librechat[bot] wants to merge 1 commit into
lia-by-librechat[bot] wants to merge 1 commit into
Conversation
Contributor
Author
|
Ready for review at the exact pushed head: This head adds an isolated context-selection experiment, independently authored synthetic cases, opt-in provider adapters and a reproducible offline-results record. Production behavior is unchanged. Actually rerun: 52 selection/HTTP-emulated tests passed; default, tight-budget and saved-score replays passed; production-import isolation, Black 24.4.0, compilation, dependency consistency and diff checks passed. Live embeddings/Jev, provider calibration, representative corpus/answer-quality evaluation, full existing RAG unit tests and database/extraction integration have not been run. No live-model performance claim is made. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add an independently authored, isolated experiment for answer-evidence selection. Compare fixed candidate pools using original order, BM25, embedding similarity and Jev usefulness scores. No production retrieval, parsing, authorization, ingestion or storage behavior changes.
Design
Checks actually run
Rerun locally on 2026-09-30 with Python 3.12.3:
pip check, diff checksThe corpus has 10 synthetic cases and 82 passages, covering exact identifiers, exceptions, negation, conflicting sources, multilingual text, duplicates, late evidence and empty/irrelevant pools. The tests emulate HTTP responses; they do not measure live provider quality.
evals/context_selection/RESULTS.mdrecords the measured offline outcomes and corpus fingerprint. Lexical filtering loses required evidence in some cases. Shorter context is not assumed to be an improvement, and no production strategy is selected from these synthetic results.Not run
Live embeddings or Jev calls, provider calibration, representative authorized LibreChat corpus replay, generated-answer evaluation, real-provider latency/cost measurements, full existing RAG unit suite, and database/extraction integration. Credentials are not available yet. Embeddings and Jev are explicitly reported as skipped, not simulated quality results. No TypeScript changes, so no TypeScript typecheck applies.
Rollout
This PR changes only
evals/and the focused test module. The existing CI unit-test run discovers that module. No dependencies, routes, schemas or production configuration change.The README documents the future credential configuration and the acceptance gates for any opt-in production integration. Authorization must precede inference egress, citation provenance must survive selection, and query-specific filtering must never replace whole-file content inspection.