Skip to content

馃И test: Evaluate Context Selection With Offline Replay - #332

Open
lia-by-librechat[bot] wants to merge 1 commit into
mainfrom
lia/context-selection-bench
Open

lia-by-librechat[bot] wants to merge 1 commit into
mainfrom
lia/context-selection-bench

Conversation

@lia-by-librechat

Copy link
Copy Markdown
Contributor

Summary

Add an independently authored, isolated experiment for answer-evidence selection. Compare fixed candidate pools using original order, BM25, embedding similarity and Jev usefulness scores. No production retrieval, parsing, authorization, ingestion or storage behavior changes.

fixed, already-authorized passages + query
  -> score once
  -> top-k versus threshold selection over the same scores
  -> original passage IDs/sources + retained and missing evidence metrics

Design

  • Preserve complete passage text and citation identity. A result ceiling is not an instruction to fill every slot.
  • Distinguish an honest empty evidence set from provider failure.
  • Track evidence precision and recall, lost facts and exceptions, correct abstentions, selected characters, scoring time and provider-reported usage.
  • Replay saved scores to sweep budgets/thresholds without repeating inference.
  • Require explicit network opt-in and provider configuration. Default runs make no provider requests, even if credentials are present.
  • Bound provider concurrency and deadlines; drain sibling tasks on cancellation/failure. Validate response indices, finite scores and embedding dimensions. Do not log response bodies or credentials.
  • Keep the evaluation independent of application initialization and database/extraction dependencies.

Checks actually run

Rerun locally on 2026-09-30 with Python 3.12.3:

Check Result
Selection and HTTP-emulated adapter tests 52 passed, zero failures/errors/skips
Offline default-budget replay Passed
Offline two-passage-budget replay Passed
Saved-score replay at four-passage budget Passed, no new provider calls
Production-import isolation Passed
Black 24.4.0, compilation, pip check, diff checks Passed
python -m pytest --noconftest tests/test_context_selection.py --junitxml=.venv/tests.xml
python3 -m evals.context_selection.replay --output .venv/default.json
python3 -m evals.context_selection.replay --max-results 2 --output .venv/tight.json
python3 -m evals.context_selection.replay --max-results 4 --reuse-scores .venv/tight.json --output .venv/reused.json

The corpus has 10 synthetic cases and 82 passages, covering exact identifiers, exceptions, negation, conflicting sources, multilingual text, duplicates, late evidence and empty/irrelevant pools. The tests emulate HTTP responses; they do not measure live provider quality.

evals/context_selection/RESULTS.md records the measured offline outcomes and corpus fingerprint. Lexical filtering loses required evidence in some cases. Shorter context is not assumed to be an improvement, and no production strategy is selected from these synthetic results.

Not run

Live embeddings or Jev calls, provider calibration, representative authorized LibreChat corpus replay, generated-answer evaluation, real-provider latency/cost measurements, full existing RAG unit suite, and database/extraction integration. Credentials are not available yet. Embeddings and Jev are explicitly reported as skipped, not simulated quality results. No TypeScript changes, so no TypeScript typecheck applies.

Rollout

This PR changes only evals/ and the focused test module. The existing CI unit-test run discovers that module. No dependencies, routes, schemas or production configuration change.

The README documents the future credential configuration and the acceptance gates for any opt-in production integration. Authorization must precede inference egress, citation provenance must survive selection, and query-specific filtering must never replace whole-file content inspection.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Ready for review at the exact pushed head:

da076b2baab398b71b7f4f0015bc9a42147ca7d1

This head adds an isolated context-selection experiment, independently authored synthetic cases, opt-in provider adapters and a reproducible offline-results record. Production behavior is unchanged.

Actually rerun: 52 selection/HTTP-emulated tests passed; default, tight-budget and saved-score replays passed; production-import isolation, Black 24.4.0, compilation, dependency consistency and diff checks passed.

Live embeddings/Jev, provider calibration, representative corpus/answer-quality evaluation, full existing RAG unit tests and database/extraction integration have not been run. No live-model performance claim is made.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant