[benchmark]: Add multi-retriever benchmarking, Nova 2 Lite support, and PGA bio/stat split - #416
[benchmark]: Add multi-retriever benchmarking, Nova 2 Lite support, and PGA bio/stat split#416oussamahansal wants to merge 4 commits into
Conversation
|
Lexical Graph Coverage Report: The coverage is at 61.9% (target: 80%). Download the HTML report here. |
mykola-pereyma
left a comment
There was a problem hiding this comment.
A few questions:
-
Three concerns bundled in one PR — The Nova 2 Lite null fix is a library-level change, the benchmark infra is test tooling, and the boto3 range is dependency management. Would it be possible to split into separate PRs? Makes review/revert cleaner if something regresses.
-
See inline comment on
on-start.shabout the boto3 version range width. -
See inline comment on
retriever_factory.pyabout themax_search_resultsscope change. -
batch_topic_extractor_sync.pyis missing a newline at end of file — minor but may trigger linting.
|
Lexical Graph Coverage Report: The coverage is at 61.9% (target: 80%). Download the HTML report here. |
|
Lexical Graph Coverage Report: The coverage is at 61.9% (target: 80%). Download the HTML report here. |
Description
Adds multi-retriever benchmark infrastructure, fixes Nova 2 Lite extraction compatibility,
and introduces PGA bio/stat split evaluation.
Changes
Problem
Related issue (if any): #
Testing
pytest)Checklist
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.