This repo holds the dataset used for contextual benchmarking of LLM-based judgement for evaluating user requests and system responses.
The dataset contains 100 user requests, with 6 system reponses per request. Five responses are incorrect and include e.g., time errors in the recommendation, while one response is correct.