Skip to content

Repository files navigation

Self-Consistency vs. Source-Grounded NLI for Hallucination Detection

Empirical comparison of two black-box hallucination-detection strategies for LLM-generated summaries, measuring where behavioral self-consistency checking succeeds and fails relative to direct source-grounded entailment.


1. Key Findings

Research Question

For claims injected into LLM-generated summaries:

How does the catch rate of black-box self-consistency checking (Method A) vary across claim types, and does source-grounded entailment (Method B) catch a distinct set of failures?

Model under test: openai/gpt-4o-mini, accessed through the OpenRouter API for both generation and verification.

Catch Rate (Recall) on Injected Unsupported Claims

Category N Method A: Self-Consistency Method B: Source-Grounded NLI
Numerical fabrication 11 100.0% (11/11) 100.0% (11/11)
Entity fabrication 10 90.0% (9/10) 100.0% (10/10)
Causal overreach 25 96.0% (24/25) 100.0% (25/25)
Overall 46 95.7% (44/46) 100.0% (46/46)

False Positive Rate on the 86 Supported Controls

Method False Positives FPR
Method A: Self-Consistency 5 / 86 5.8%
Method B: Source-Grounded NLI 4 / 86 4.7%

Cross-Tabulation: Unsupported Claims Only

Method B Caught Method B Missed
Method A Caught 44 0
Method A Missed 2 0

Takeaway

Method B strictly dominates on recall: it catches 100% of injected unsupported claims, including the two claims missed by Method A.

Method A, despite having no access to the source document during verification, achieves 95.7% recall. Its failures occur primarily when an LLM reproduces a plausible narrative schema across independently generated summaries and the equivalence verifier mistakes this consistency for factual corroboration.

The results highlight a fundamental distinction:

Self-consistency measures consensus; source-grounded NLI measures textual support.


2. Repository Architecture

.
├── .env.example                 # Template for environment variables (API keys)
├── .gitignore                   # Ignores .env, .venv, caches, and IDE configs
├── README.md                    # Project overview, methodology, results, and reproduction guide
├── research_report.md           # Detailed research write-up and error analysis
│
├── claims_dataset.py            # Defines the 132-claim adversarial benchmark (13 source docs)
├── claims_dataset.json          # Benchmark dataset JSON (generated by claims_dataset.py)
│
├── method_a_self_consistency.py # Method A: N=5 regeneration + agreement scoring
├── method_b_entailment.py       # Method B: single-pass source-grounded NLI
│
├── evaluate_and_score.py        # Calculates metrics and cross-tabulation
├── check_missed_claims.py       # Inspects Method A false negatives
│
├── regenerations_cache.json     # Cached N=5 summaries for Method A
├── method_a_results.json        # Per-claim Method A results
├── method_b_results.json        # Per-claim Method B results
└── evaluation_summary.json      # Aggregated evaluation metrics

.env, .venv/, and __pycache__/ are deliberately excluded from the repository through .gitignore.

Dataset Generation

claims_dataset.json is generated by running:

python claims_dataset.py

The downstream evaluation scripts read the serialized JSON directly from disk.


3. Methodology

3.1 Adversarial Dataset

The benchmark is defined in claims_dataset.py and contains:

  • 132 total claims
  • 13 source documents
  • 46 unsupported injected claims
  • 86 supported control claims

The source documents cover multiple technical and factual domains, including:

  • Astronomy press releases
  • Corporate earnings reports
  • Municipal continuity orders
  • Clinical trial readouts
  • Semiconductor hardware announcements

Each dataset entry contains:

claim_id
doc_id
source_text
claim_text
category
is_unsupported

Claim Categories

Unsupported Injections — N=46

Category N Description
numerical 11 Altered metrics, dates, percentages, and sample sizes
entity 10 Fabricated researchers, organizations, partners, or instruments
causal 25 Unsupported causal extrapolations and premature efficacy claims

Supported Controls — N=86

Grounded factual claims and appropriately hedged statements where:

is_unsupported = False

3.2 Method A — Self-Consistency

Implementation:

method_a_self_consistency.py

Method A is deliberately source-blind at verification time. It attempts to detect hallucinations by measuring whether an injected claim consistently appears across independently generated summaries.

Procedure

  1. For each of the 13 source documents, generate 5 alternative summaries.
  2. Generation uses:
N_REGENERATIONS = 5
temperature = 0.7
  1. Generated summaries are cached in:
regenerations_cache.json
  1. For each of the 132 claims, run a binary equivalence check against each of the 5 summaries.
  2. Each equivalence check uses:
temperature = 0.0
  1. Calculate the agreement score:
agreement_score = matches / 5
  1. Flag the claim as unsupported when:
agreement_score <= 0.40

In other words, a claim is considered "caught" when it appears to be supported by no more than 2 of the 5 generated summaries.

Important Property

Method A does not inspect the original source text during verification.

Its signal is therefore behavioral consistency rather than direct factual grounding.


3.3 Method B — Source-Grounded NLI

Implementation:

method_b_entailment.py

Method B performs a direct source-grounded entailment check.

For each claim, the verifier receives the relevant source text and determines whether the source directly supports the claim.

Procedure

For each claim:

  1. Provide the source text.
  2. Provide the claim.
  3. Ask whether the source directly supports the assertion.
  4. Run the verification prompt at:
temperature = 0.0
  1. Parse the categorical output:
VERDICT: YES

or:

VERDICT: NO
  1. A NO verdict marks the unsupported claim as caught.

Method B therefore has direct access to the evidence against which the claim is evaluated.


3.4 Scoring

Evaluation is performed by:

evaluate_and_score.py

The script joins:

method_a_results.json
method_b_results.json

using:

claim_id

It then calculates:

  • Overall recall
  • Category-level recall
  • False positive rate on supported controls
  • Cross-tabulation between Method A and Method B

Qualitative inspection of Method A's false negatives is performed by:

check_missed_claims.py

4. Detailed Results & Failure Analysis

Method A missed exactly 2 of the 46 injected unsupported claims. Both were correctly caught by Method B.

Failure Case 1 — Entity Fabrication

Claim ID: doc11_c04 Category: Entity Fabrication Method A Agreement Score: 0.8 / 80%

Injected Claim

"City Attorney Jennifer Martinez advised the council that state law allows local age restrictions on e-bike operators."

Source Truth

The source states that councilmembers asked municipal staff to pursue legislative options. No city attorney named Jennifer Martinez was identified as having advised the council.

Why Method A Failed

The generated summaries repeatedly focused on the municipal regulatory debate. The equivalence verifier therefore encountered a plausible narrative structure involving:

  • Municipal regulation
  • City officials
  • State law
  • E-bike restrictions

Because these concepts appeared consistently across generations, the verifier treated the fabricated official attribution as semantically supported.

Failure Mechanism

The verifier exhibited semantic drift:

Topical consistency
        ↓
Similar narrative framing
        ↓
Perceived semantic equivalence
        ↓
False acceptance

The problem is that contextual plausibility does not establish factual support.


Failure Case 2 — Causal Overreach

Claim ID: doc13_c06 Category: Causal Overreach Method A Agreement Score: 1.0 / 100%

Injected Claim

"Webb's coronagraph technology enabled this discovery, which would have been completely impossible with any other existing telescope."

Source Truth

The source describes how the coronagraph operates and its role in the observation. However, it does not establish that:

  • the discovery was exclusively enabled by Webb's coronagraph, or
  • the discovery would have been completely impossible using every other existing telescope.

Why Method A Failed

Every generated summary emphasized the importance of the coronagraph. The equivalence verifier therefore observed extremely high agreement:

5 / 5 summaries
=
agreement score of 1.0

However, the injected claim contained a much stronger assertion than the source:

"The coronagraph was useful"

became:

"The coronagraph made the discovery uniquely possible."

Failure Mechanism

Method A conflated repeated rhetorical emphasis with evidence for an absolute causal claim.

This is particularly problematic for claims containing:

  • "only"
  • "completely impossible"
  • "the first"
  • "the only"
  • "always"
  • "never"
  • "caused"
  • "enabled"
  • "proved"

5. What the Failure Cases Demonstrate

The two false negatives expose the central theoretical limitation of self-consistency-based hallucination detection:

A model can consistently repeat a hallucination.

Independent generations are not independent sources of truth. They are samples from the same model's learned distribution.

If the model has a strong prior about a particular narrative structure, multiple generations can converge on the same unsupported assertion.

Therefore:

High agreement
    ≠
High factuality

More precisely:

Self-consistency
    measures
behavioral agreement

Source-grounded NLI
    measures
evidence support

This distinction explains why Method B was able to catch both of Method A's false negatives.


6. Quickstart

Requirements

  • Python 3.10 or higher
  • An OpenRouter API key (or OpenAI-compatible provider)

6.1 Clone the Repository

git clone <YOUR_REPOSITORY_URL>
cd <YOUR_REPOSITORY_DIRECTORY>

6.2 Create and Activate a Virtual Environment

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

Windows

python -m venv .venv
.venv\Scripts\activate

6.3 Install Dependencies

pip install openai tqdm python-dotenv

6.4 Configure API Credentials

Copy the example environment file:

cp .env.example .env

Then edit .env:

OPENROUTER_API_KEY=your_key_here

Do not commit .env to GitHub. The repository's .gitignore is configured to exclude it.


7. Reproduction Pipeline

Run the scripts sequentially.

Step 1 — Generate Dataset

python claims_dataset.py

This creates:

claims_dataset.json

Step 2 — Run Method A

python method_a_self_consistency.py

This generates:

regenerations_cache.json
method_a_results.json

The first file caches the five generated summaries per source document.

Step 3 — Run Method B

python method_b_entailment.py

This generates:

method_b_results.json

Step 4 — Evaluate Both Methods

python evaluate_and_score.py

This generates:

evaluation_summary.json

The evaluation contains:

  • Category-level recall
  • Overall recall
  • False positive rates
  • Cross-tabulation

Step 5 — Inspect Method A Misses

python check_missed_claims.py

This prints the unsupported claims that:

  • Method A missed
  • Method B caught

This provides a qualitative error analysis of the self-consistency approach.


8. Caching

method_a_self_consistency.py caches generated summaries in:

regenerations_cache.json

This prevents unnecessary regeneration when rerunning the evaluation.

To regenerate all summaries from scratch, delete the cache:

rm regenerations_cache.json

Then rerun:

python method_a_self_consistency.py

9. Environment Variables

The project expects:

OPENROUTER_API_KEY=your_key_here

The scripts use python-dotenv to load the value from .env.

Example .env.example:

OPENROUTER_API_KEY=

10. Reproducibility Notes

The benchmark is deterministic at the verification stage because both verification methods use:

temperature = 0.0

Method A's summary generation intentionally uses stochastic generation:

temperature = 0.7
N = 5

Therefore, deleting regenerations_cache.json and regenerating summaries can potentially produce different intermediate generations and agreement scores.

For reproducing the reported results as closely as possible, use the committed:

regenerations_cache.json
method_a_results.json
method_b_results.json
evaluation_summary.json

when available.


11. Threats to Validity

11.1 Single-Model Evaluator

Both summary generation and verification rely on:

openai/gpt-4o-mini

Therefore, the results do not establish that the same performance would occur with other models.

Cross-model verification using models such as Claude, Gemini, or Llama was not evaluated.

11.2 Dataset Scale

The benchmark contains:

132 claims
13 source documents
46 unsupported claims
86 supported controls

The dataset was constructed by a single author.

No multi-annotator inter-rater reliability analysis, such as Cohen's Kappa, was performed.

11.3 Threshold Parameterization

Method A uses:

N = 5 regenerations
threshold <= 0.40

This represents an explicit design choice.

Changing either parameter could alter the precision/recall trade-off.

For example:

More regenerations
        ↓
Higher computational cost
        ↓
Potentially more stable agreement estimates

Changing the threshold can also alter how aggressively low-consensus claims are flagged.

11.4 Prompt Dependence

Both methods depend on the exact prompts used for:

  • Summary generation
  • Equivalence verification
  • Source-grounded entailment

Small changes to prompt wording may affect model behavior.

11.5 Synthetic Adversarial Claims

The benchmark consists of deliberately injected unsupported claims.

Although the categories are designed to represent common hallucination patterns, performance on naturally occurring hallucinations in production systems may differ.


12. Interpretation

The benchmark suggests that source-grounded verification provides a stronger signal for detecting unsupported claims than self-consistency alone.

Method A

Advantages:

  • Does not require source access during verification
  • Can detect claims that fail to recur across generations
  • Achieves high recall despite being source-blind
  • Can potentially be applied when only generated outputs are available

Limitations:

  • Consensus can reinforce the same hallucination
  • Vulnerable to plausible entity fabrication
  • Vulnerable to causal overreach
  • Requires multiple generations
  • Higher inference cost due to repeated generation and verification

Method B

Advantages:

  • Directly evaluates claims against source evidence
  • Achieved 100% recall on this benchmark
  • Caught both Method A false negatives
  • Requires only a single verification pass per claim

Limitations:

  • Requires access to the source document
  • Depends on the verifier's ability to interpret the source correctly
  • Still produced false positives on 4 supported controls
  • Remains dependent on model and prompt quality

13. Conclusion

This benchmark provides an empirical comparison between two fundamentally different hallucination-detection strategies.

Method A — Self-Consistency

Claim
  ↓
Compare against multiple generated summaries
  ↓
Measure recurrence
  ↓
Flag low-consensus claims

Method B — Source-Grounded NLI

Claim
  ↓
Compare directly against source
  ↓
Determine entailment
  ↓
Flag unsupported claims

On the 46 injected unsupported claims:

Method A: 95.7% recall
Method B: 100.0% recall

On the 86 supported controls:

Method A: 5.8% false positive rate
Method B: 4.7% false positive rate

The two Method A failures demonstrate that consistency is not equivalent to truth.

An LLM can repeatedly produce the same unsupported entity attribution or causal interpretation because the claim fits a plausible narrative pattern.

The results therefore support a practical distinction:

Self-consistency can be a useful hallucination signal, but source-grounded verification provides a stronger basis for determining whether a claim is actually supported by evidence.


14. Project Structure at a Glance

                 ┌──────────────────────┐
                 │   Source Documents   │
                 └──────────┬───────────┘
                            │
                            ▼
                 ┌──────────────────────┐
                 │  Adversarial Claims  │
                 │       Dataset        │
                 └──────────┬───────────┘
                            │
              ┌─────────────┴─────────────┐
              │                           │
              ▼                           ▼
    ┌───────────────────┐       ┌────────────────────┐
    │     Method A      │       │      Method B       │
    │ Self-Consistency   │       │ Source-Grounded NLI│
    └─────────┬─────────┘       └──────────┬─────────┘
              │                            │
              ▼                            ▼
    ┌───────────────────┐       ┌────────────────────┐
    │ method_a_results   │       │ method_b_results   │
    │       .json        │       │       .json        │
    └─────────┬─────────┘       └──────────┬─────────┘
              │                            │
              └─────────────┬──────────────┘
                            ▼
                 ┌──────────────────────┐
                 │   Evaluation &       │
                 │   Error Analysis     │
                 └──────────┬───────────┘
                            │
                            ▼
                 ┌──────────────────────┐
                 │ evaluation_summary   │
                 │        .json         │
                 └──────────────────────┘

License

MIT License. See LICENSE for details.

Citation

If you use this benchmark or methodology in research, please cite this repository.

About

Adversarial evaluation comparing black-box self-consistency checking vs. source-grounded NLI for hallucination detection in LLM summaries.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages