Skip to content
@JUDGE-Bench

JUDGE-Bench

Benchmarking Contextual Understanding for In-Car Conversational Systems

This repo hold the dataset and evaluation code of the paper “Benchmarking Contextual Understanding for In-Car Conversational Systems”. The code is used to analyse LLM-based judgement for different large language models and various prompting techniques.

Dataset

The dataset can be found here: dataset

The dataset is licensed under the accompanied license.

Code

The evaluation code can be accessed via the following link: code

The code is licensed under the accompanied license.

Leaderboard

The overall performance of LLM-based judgement techniques and models on the benchmarking dataset is provided below:

More results you can find in the paper.

When using the dataset our refernce to our results please cite the paper as follows.

@misc{habicht2025benchmarking,
      title={Benchmarking Contextual Understanding for In-Car Conversational Systems}, 
      author={Philipp Habicht and Lev Sorokin and Abdullah Saydemir and Ken E. Friedl and Andrea Stocco},
      year={2025},
      eprint={2512.12042},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2512.12042}, 
}

Popular repositories Loading

  1. dataset dataset Public

    1

  2. .github .github Public

  3. code code Public

    Code for benchmarking different LLM with different prompting strategies.

    Python

Repositories

Showing 3 of 3 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…