Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Dataset for Contextual Benchmarking of LLM-based Judgement

This repo holds the dataset used for contextual benchmarking of LLM-based judgement for evaluating user requests and system responses.

The dataset contains 100 user requests, with 6 system reponses per request. Five responses are incorrect and include e.g., time errors in the recommendation, while one response is correct.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors