Repository navigation
feat(mem0): eval sets by Bob subagents + local reranker run - #22
Conversation
…13, rollback demo)
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Downshift cost diffNo LLM call site cost changes in Projection, not a bill: tokens estimated from source (words in resolved prompt text; output = max_tokens, or 256 if unset) x illustrative prices x configured calls/day x 30 days. |
…, re-run reranker
|
Bob built-in code review (/review, 34 tools, 1.98 coins) found 3 issues, all fixed by hand: rr-11 and rr-19 had wrong expected bands (documents that directly answer the query), and the report printed 'No payback' when payback was never computed. Only the 2 relabeled cases were re-run. Final: 7b 17/24, 3b 13/24, 1.5b 13/24, 0.5b 11/24; best candidates keep 76% of baseline, no safe downgrade. |
Bob B12 wrote 51 eval cases for 3 mem0 features with 3 parallel subagents. Local reranker run on Qwen tiers: 7b 71%, 3b 62%, 1.5b 50%, 0.5b 50%, so no safe downgrade (3b keeps 88% < 95%); failures cluster in the partial-relevance band. Analysis cost $0.29. Report fix: clear sentence when there is no payback. Bob B13 updated the case-study doc and web (one invented claim corrected by hand); a README edit was undone with Bob's rollback.