test Llm/deepeval api tool calling - #232
Conversation
RAG System Evaluation ReportDeepEval Test Results Summary
Total Tests: 6 | Passed: 0 | Failed: 6 Detailed Test Results| Test | Language | Category | CP | CR | CRel | AR | Faith | Status | Legend: CP = Contextual Precision, CR = Contextual Recall, CRel = Contextual Relevancy, AR = Answer Relevancy, Faith = Faithfulness Failed Test Analysis
RecommendationsContextual Precision (Score: 0.000): Consider improving your reranking model or adjusting reranking parameters to better prioritize relevant documents. Contextual Recall (Score: 0.000): Review your embedding model choice and vector search parameters. Consider domain-specific embeddings. Contextual Relevancy (Score: 0.000): Optimize chunk size and top-K retrieval parameters to reduce noise in retrieved contexts. Report generated on 2026-06-26 03:52:40 by DeepEval automated testing pipeline |
API Tool Calling Evaluation ReportIssue buerokratt#447 — DeepEval coverage for the API Tool Calling feature. Started:
Results by scenario type
Detailed resultsNo scenarios recorded. Methodology
|
No description provided.