Skip to content

test Llm/deepeval api tool calling - #232

Open
nuwangeek wants to merge 5 commits into
deepeval-tempfrom
llm/deepeval-api-tool-calling
Open

test Llm/deepeval api tool calling#232
nuwangeek wants to merge 5 commits into
deepeval-tempfrom
llm/deepeval-api-tool-calling

Conversation

@nuwangeek

Copy link
Copy Markdown

No description provided.

@github-actions

github-actions Bot commented Jun 25, 2026

Copy link
Copy Markdown

RAG System Evaluation Report

DeepEval Test Results Summary

Metric Pass Rate Avg Score Status
Overall 0.0% - FAIL
Contextual Precision 0.0% 0.000 FAIL
Contextual Recall 0.0% 0.000 FAIL
Contextual Relevancy 0.0% 0.000 FAIL
Answer Relevancy 83.3% 0.833 PASS
Faithfulness 100.0% 1.000 PASS

Total Tests: 6 | Passed: 0 | Failed: 6
Test Duration: 11.2 minutes

Detailed Test Results

| Test | Language | Category | CP | CR | CRel | AR | Faith | Status |
|------|----------|----------|----|----|------|----|----- -|--------|
| 1 | ET | mobile_id_usage | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | FAIL |
| 2 | ET | digital_identity_security | 0.00 | 0.00 | 0.00 | 0.00 | 1.00 | FAIL |
| 3 | ET | digital_identity | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | FAIL |
| 4 | EN | digital_identity | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | FAIL |
| 5 | ET | digital_identity | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | FAIL |
| 6 | ET | statistics | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | FAIL |

Legend: CP = Contextual Precision, CR = Contextual Recall, CRel = Contextual Relevancy, AR = Answer Relevancy, Faith = Faithfulness
Languages: EN = English, ET = Estonian, RU = Russian

Failed Test Analysis

Test Query Metric Score Issue
1 Mida teha kui mobiil-ID kasutamisel kinnituskood e... contextual_precision 0.00 The score is 0.00 because there are no nodes in the retrieval contexts, so there is no opportunity for relevant information to be ranked above irrelevant information.
1 Mida teha kui mobiil-ID kasutamisel kinnituskood e... contextual_recall 0.00 The score is 0.00 because none of the sentences in the expected output can be attributed to any node(s) in retrieval context, as the retrieval context is empty (0 nodes).
1 Mida teha kui mobiil-ID kasutamisel kinnituskood e... contextual_relevancy 0.00 The score is 0.00 because there are no relevant statements in the retrieval context and no reasons for irrelevancy were provided.
2 Mida teha, kui minu telefon Mobiil-ID-ga varastata... contextual_precision 0.00 The score is 0.00 because there are no nodes in the retrieval contexts, so there is no opportunity for relevant information to be ranked above irrelevant information.
2 Mida teha, kui minu telefon Mobiil-ID-ga varastata... contextual_recall 0.00 The score is 0.00 because none of the sentences in the expected output can be attributed to any node(s) in retrieval context, as the retrieval context is empty.
2 Mida teha, kui minu telefon Mobiil-ID-ga varastata... contextual_relevancy 0.00 The score is 0.00 because there are no relevant statements in the retrieval context and no reasons provided for irrelevancy, indicating a complete lack of connection to the input.
2 Mida teha, kui minu telefon Mobiil-ID-ga varastata... answer_relevancy 0.00 The score is 0.00 because the output contains only irrelevant statements, such as apologies, requests for clarification, and comments on context, without providing any answer to the user's question about what to do if their phone with Mobile-ID is stolen.
3 Mis on eIDAS määrus? contextual_precision 0.00 The score is 0.00 because there are no nodes in the retrieval contexts, so there is no opportunity for relevant nodes to be ranked higher than irrelevant ones.
3 Mis on eIDAS määrus? contextual_recall 0.00 The score is 0.00 because none of the sentences in the expected output can be attributed to any node(s) in retrieval context, as the retrieval context is empty.
3 Mis on eIDAS määrus? contextual_relevancy 0.00 The score is 0.00 because there are no relevant statements in the retrieval context and no reasons for irrelevancy were provided.
4 Why am I getting an error when trying to sign docu... contextual_precision 0.00 The score is 0.00 because there are no nodes in the retrieval contexts, so no relevant information is ranked above irrelevant information. The absence of any nodes means the system did not retrieve any content to evaluate.
4 Why am I getting an error when trying to sign docu... contextual_recall 0.00 The score is 0.00 because none of the sentences in the expected output can be attributed to any node(s) in the retrieval context, as there are no relevant nodes present.
4 Why am I getting an error when trying to sign docu... contextual_relevancy 0.00 The score is 0.00 because there are no relevant statements or reasons provided in the retrieval context that address the input question about DigiDoc4 errors.
5 Kuidas aktiveerida Mobiil-ID? contextual_precision 0.00 The score is 0.00 because there are no nodes in the retrieval contexts, so there is no opportunity for relevant information to be ranked higher than irrelevant information.
5 Kuidas aktiveerida Mobiil-ID? contextual_recall 0.00 The score is 0.00 because none of the sentences in the expected output can be linked to any node(s) in the retrieval context, as there are no nodes present.
5 Kuidas aktiveerida Mobiil-ID? contextual_relevancy 0.00 The score is 0.00 because there are no relevant statements in the retrieval context and no reasons for irrelevancy were provided.
6 Mis on Eesti sotsiaaluuring ja miks ma peaksin osa... contextual_precision 0.00 The score is 0.00 because there are no nodes in the retrieval contexts, so no relevant information is ranked above irrelevant information. The absence of any nodes means the system did not retrieve any content to evaluate.
6 Mis on Eesti sotsiaaluuring ja miks ma peaksin osa... contextual_recall 0.00 The score is 0.00 because none of the sentences in the expected output can be attributed to any node(s) in retrieval context, as the retrieval context is empty.
6 Mis on Eesti sotsiaaluuring ja miks ma peaksin osa... contextual_relevancy 0.00 The score is 0.00 because there are no relevant statements in the retrieval context and no reasons provided for irrelevancy, indicating a complete lack of connection to the input.

Recommendations

Contextual Precision (Score: 0.000): Consider improving your reranking model or adjusting reranking parameters to better prioritize relevant documents.

Contextual Recall (Score: 0.000): Review your embedding model choice and vector search parameters. Consider domain-specific embeddings.

Contextual Relevancy (Score: 0.000): Optimize chunk size and top-K retrieval parameters to reduce noise in retrieved contexts.


Report generated on 2026-06-26 03:52:40 by DeepEval automated testing pipeline

@github-actions

github-actions Bot commented Jun 25, 2026

Copy link
Copy Markdown

API Tool Calling Evaluation Report

Issue buerokratt#447 — DeepEval coverage for the API Tool Calling feature.

Started: 2026-06-26T00:50:12.336504

Metric Value
Total scenarios 0
Passed 0
Failed 0
Errored 0
Pass rate 0.0%

Results by scenario type

Type Metric Passed Total Pass rate

Detailed results

No scenarios recorded.

Methodology

  • Strict scenarios (issue specifies the expected endpoint and params) are scored with ToolCorrectnessMetric, evaluation_params=[ToolCallParams.INPUT_PARAMETERS], threshold 1.0. Tool name must match exactly and every expected parameter must be present with the expected value (extra parameters allowed).
  • Loose scenarios (issue describes the flow but not the expected resolution) are scored with ArgumentCorrectnessMetric (LLM-as-judge, threshold 0.7).
  • Multi-intent scenarios use ArgumentCorrectnessMetric when the agent resolves a tool call; if the agent asks a clarifying question instead, only routing-to-ATC is verified (rules out silent fall-through to RAG/OOD).
  • API tool endpoints are seeded into the testcontainers-backed Qdrant from tests/api_tool_eval/test-endpoints.json via the api_tool_endpoints_indexed fixture in conftest.py.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant