Empirical comparison of two black-box hallucination-detection strategies for LLM-generated summaries, measuring where behavioral self-consistency checking succeeds and fails relative to direct source-grounded entailment.
For claims injected into LLM-generated summaries:
How does the catch rate of black-box self-consistency checking (Method A) vary across claim types, and does source-grounded entailment (Method B) catch a distinct set of failures?
Model under test: openai/gpt-4o-mini, accessed through the OpenRouter API for both generation and verification.
| Category | N | Method A: Self-Consistency | Method B: Source-Grounded NLI |
|---|---|---|---|
| Numerical fabrication | 11 | 100.0% (11/11) | 100.0% (11/11) |
| Entity fabrication | 10 | 90.0% (9/10) | 100.0% (10/10) |
| Causal overreach | 25 | 96.0% (24/25) | 100.0% (25/25) |
| Overall | 46 | 95.7% (44/46) | 100.0% (46/46) |
| Method | False Positives | FPR |
|---|---|---|
| Method A: Self-Consistency | 5 / 86 | 5.8% |
| Method B: Source-Grounded NLI | 4 / 86 | 4.7% |
| Method B Caught | Method B Missed | |
|---|---|---|
| Method A Caught | 44 | 0 |
| Method A Missed | 2 | 0 |
Method B strictly dominates on recall: it catches 100% of injected unsupported claims, including the two claims missed by Method A.
Method A, despite having no access to the source document during verification, achieves 95.7% recall. Its failures occur primarily when an LLM reproduces a plausible narrative schema across independently generated summaries and the equivalence verifier mistakes this consistency for factual corroboration.
The results highlight a fundamental distinction:
Self-consistency measures consensus; source-grounded NLI measures textual support.
.
├── .env.example # Template for environment variables (API keys)
├── .gitignore # Ignores .env, .venv, caches, and IDE configs
├── README.md # Project overview, methodology, results, and reproduction guide
├── research_report.md # Detailed research write-up and error analysis
│
├── claims_dataset.py # Defines the 132-claim adversarial benchmark (13 source docs)
├── claims_dataset.json # Benchmark dataset JSON (generated by claims_dataset.py)
│
├── method_a_self_consistency.py # Method A: N=5 regeneration + agreement scoring
├── method_b_entailment.py # Method B: single-pass source-grounded NLI
│
├── evaluate_and_score.py # Calculates metrics and cross-tabulation
├── check_missed_claims.py # Inspects Method A false negatives
│
├── regenerations_cache.json # Cached N=5 summaries for Method A
├── method_a_results.json # Per-claim Method A results
├── method_b_results.json # Per-claim Method B results
└── evaluation_summary.json # Aggregated evaluation metrics
.env, .venv/, and __pycache__/ are deliberately excluded from the repository through .gitignore.
claims_dataset.json is generated by running:
python claims_dataset.pyThe downstream evaluation scripts read the serialized JSON directly from disk.
The benchmark is defined in claims_dataset.py and contains:
- 132 total claims
- 13 source documents
- 46 unsupported injected claims
- 86 supported control claims
The source documents cover multiple technical and factual domains, including:
- Astronomy press releases
- Corporate earnings reports
- Municipal continuity orders
- Clinical trial readouts
- Semiconductor hardware announcements
Each dataset entry contains:
claim_id
doc_id
source_text
claim_text
category
is_unsupported
| Category | N | Description |
|---|---|---|
numerical |
11 | Altered metrics, dates, percentages, and sample sizes |
entity |
10 | Fabricated researchers, organizations, partners, or instruments |
causal |
25 | Unsupported causal extrapolations and premature efficacy claims |
Grounded factual claims and appropriately hedged statements where:
is_unsupported = False
Implementation:
method_a_self_consistency.py
Method A is deliberately source-blind at verification time. It attempts to detect hallucinations by measuring whether an injected claim consistently appears across independently generated summaries.
- For each of the 13 source documents, generate 5 alternative summaries.
- Generation uses:
N_REGENERATIONS = 5
temperature = 0.7
- Generated summaries are cached in:
regenerations_cache.json
- For each of the 132 claims, run a binary equivalence check against each of the 5 summaries.
- Each equivalence check uses:
temperature = 0.0
- Calculate the agreement score:
agreement_score = matches / 5
- Flag the claim as unsupported when:
agreement_score <= 0.40
In other words, a claim is considered "caught" when it appears to be supported by no more than 2 of the 5 generated summaries.
Method A does not inspect the original source text during verification.
Its signal is therefore behavioral consistency rather than direct factual grounding.
Implementation:
method_b_entailment.py
Method B performs a direct source-grounded entailment check.
For each claim, the verifier receives the relevant source text and determines whether the source directly supports the claim.
For each claim:
- Provide the source text.
- Provide the claim.
- Ask whether the source directly supports the assertion.
- Run the verification prompt at:
temperature = 0.0
- Parse the categorical output:
VERDICT: YES
or:
VERDICT: NO
- A
NOverdict marks the unsupported claim as caught.
Method B therefore has direct access to the evidence against which the claim is evaluated.
Evaluation is performed by:
evaluate_and_score.py
The script joins:
method_a_results.json
method_b_results.json
using:
claim_id
It then calculates:
- Overall recall
- Category-level recall
- False positive rate on supported controls
- Cross-tabulation between Method A and Method B
Qualitative inspection of Method A's false negatives is performed by:
check_missed_claims.py
Method A missed exactly 2 of the 46 injected unsupported claims. Both were correctly caught by Method B.
Claim ID: doc11_c04
Category: Entity Fabrication
Method A Agreement Score: 0.8 / 80%
"City Attorney Jennifer Martinez advised the council that state law allows local age restrictions on e-bike operators."
The source states that councilmembers asked municipal staff to pursue legislative options. No city attorney named Jennifer Martinez was identified as having advised the council.
The generated summaries repeatedly focused on the municipal regulatory debate. The equivalence verifier therefore encountered a plausible narrative structure involving:
- Municipal regulation
- City officials
- State law
- E-bike restrictions
Because these concepts appeared consistently across generations, the verifier treated the fabricated official attribution as semantically supported.
The verifier exhibited semantic drift:
Topical consistency
↓
Similar narrative framing
↓
Perceived semantic equivalence
↓
False acceptance
The problem is that contextual plausibility does not establish factual support.
Claim ID: doc13_c06
Category: Causal Overreach
Method A Agreement Score: 1.0 / 100%
"Webb's coronagraph technology enabled this discovery, which would have been completely impossible with any other existing telescope."
The source describes how the coronagraph operates and its role in the observation. However, it does not establish that:
- the discovery was exclusively enabled by Webb's coronagraph, or
- the discovery would have been completely impossible using every other existing telescope.
Every generated summary emphasized the importance of the coronagraph. The equivalence verifier therefore observed extremely high agreement:
5 / 5 summaries
=
agreement score of 1.0
However, the injected claim contained a much stronger assertion than the source:
"The coronagraph was useful"
became:
"The coronagraph made the discovery uniquely possible."
Method A conflated repeated rhetorical emphasis with evidence for an absolute causal claim.
This is particularly problematic for claims containing:
- "only"
- "completely impossible"
- "the first"
- "the only"
- "always"
- "never"
- "caused"
- "enabled"
- "proved"
The two false negatives expose the central theoretical limitation of self-consistency-based hallucination detection:
A model can consistently repeat a hallucination.
Independent generations are not independent sources of truth. They are samples from the same model's learned distribution.
If the model has a strong prior about a particular narrative structure, multiple generations can converge on the same unsupported assertion.
Therefore:
High agreement
≠
High factuality
More precisely:
Self-consistency
measures
behavioral agreement
Source-grounded NLI
measures
evidence support
This distinction explains why Method B was able to catch both of Method A's false negatives.
- Python 3.10 or higher
- An OpenRouter API key (or OpenAI-compatible provider)
git clone <YOUR_REPOSITORY_URL>
cd <YOUR_REPOSITORY_DIRECTORY>python3 -m venv .venv
source .venv/bin/activatepython -m venv .venv
.venv\Scripts\activatepip install openai tqdm python-dotenvCopy the example environment file:
cp .env.example .envThen edit .env:
OPENROUTER_API_KEY=your_key_hereDo not commit .env to GitHub. The repository's .gitignore is configured to exclude it.
Run the scripts sequentially.
python claims_dataset.pyThis creates:
claims_dataset.json
python method_a_self_consistency.pyThis generates:
regenerations_cache.json
method_a_results.json
The first file caches the five generated summaries per source document.
python method_b_entailment.pyThis generates:
method_b_results.json
python evaluate_and_score.pyThis generates:
evaluation_summary.json
The evaluation contains:
- Category-level recall
- Overall recall
- False positive rates
- Cross-tabulation
python check_missed_claims.pyThis prints the unsupported claims that:
- Method A missed
- Method B caught
This provides a qualitative error analysis of the self-consistency approach.
method_a_self_consistency.py caches generated summaries in:
regenerations_cache.json
This prevents unnecessary regeneration when rerunning the evaluation.
To regenerate all summaries from scratch, delete the cache:
rm regenerations_cache.jsonThen rerun:
python method_a_self_consistency.pyThe project expects:
OPENROUTER_API_KEY=your_key_hereThe scripts use python-dotenv to load the value from .env.
Example .env.example:
OPENROUTER_API_KEY=The benchmark is deterministic at the verification stage because both verification methods use:
temperature = 0.0
Method A's summary generation intentionally uses stochastic generation:
temperature = 0.7
N = 5
Therefore, deleting regenerations_cache.json and regenerating summaries can potentially produce different intermediate generations and agreement scores.
For reproducing the reported results as closely as possible, use the committed:
regenerations_cache.json
method_a_results.json
method_b_results.json
evaluation_summary.json
when available.
Both summary generation and verification rely on:
openai/gpt-4o-mini
Therefore, the results do not establish that the same performance would occur with other models.
Cross-model verification using models such as Claude, Gemini, or Llama was not evaluated.
The benchmark contains:
132 claims
13 source documents
46 unsupported claims
86 supported controls
The dataset was constructed by a single author.
No multi-annotator inter-rater reliability analysis, such as Cohen's Kappa, was performed.
Method A uses:
N = 5 regenerations
threshold <= 0.40
This represents an explicit design choice.
Changing either parameter could alter the precision/recall trade-off.
For example:
More regenerations
↓
Higher computational cost
↓
Potentially more stable agreement estimates
Changing the threshold can also alter how aggressively low-consensus claims are flagged.
Both methods depend on the exact prompts used for:
- Summary generation
- Equivalence verification
- Source-grounded entailment
Small changes to prompt wording may affect model behavior.
The benchmark consists of deliberately injected unsupported claims.
Although the categories are designed to represent common hallucination patterns, performance on naturally occurring hallucinations in production systems may differ.
The benchmark suggests that source-grounded verification provides a stronger signal for detecting unsupported claims than self-consistency alone.
Advantages:
- Does not require source access during verification
- Can detect claims that fail to recur across generations
- Achieves high recall despite being source-blind
- Can potentially be applied when only generated outputs are available
Limitations:
- Consensus can reinforce the same hallucination
- Vulnerable to plausible entity fabrication
- Vulnerable to causal overreach
- Requires multiple generations
- Higher inference cost due to repeated generation and verification
Advantages:
- Directly evaluates claims against source evidence
- Achieved 100% recall on this benchmark
- Caught both Method A false negatives
- Requires only a single verification pass per claim
Limitations:
- Requires access to the source document
- Depends on the verifier's ability to interpret the source correctly
- Still produced false positives on 4 supported controls
- Remains dependent on model and prompt quality
This benchmark provides an empirical comparison between two fundamentally different hallucination-detection strategies.
Claim
↓
Compare against multiple generated summaries
↓
Measure recurrence
↓
Flag low-consensus claims
Claim
↓
Compare directly against source
↓
Determine entailment
↓
Flag unsupported claims
On the 46 injected unsupported claims:
Method A: 95.7% recall
Method B: 100.0% recall
On the 86 supported controls:
Method A: 5.8% false positive rate
Method B: 4.7% false positive rate
The two Method A failures demonstrate that consistency is not equivalent to truth.
An LLM can repeatedly produce the same unsupported entity attribution or causal interpretation because the claim fits a plausible narrative pattern.
The results therefore support a practical distinction:
Self-consistency can be a useful hallucination signal, but source-grounded verification provides a stronger basis for determining whether a claim is actually supported by evidence.
┌──────────────────────┐
│ Source Documents │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Adversarial Claims │
│ Dataset │
└──────────┬───────────┘
│
┌─────────────┴─────────────┐
│ │
▼ ▼
┌───────────────────┐ ┌────────────────────┐
│ Method A │ │ Method B │
│ Self-Consistency │ │ Source-Grounded NLI│
└─────────┬─────────┘ └──────────┬─────────┘
│ │
▼ ▼
┌───────────────────┐ ┌────────────────────┐
│ method_a_results │ │ method_b_results │
│ .json │ │ .json │
└─────────┬─────────┘ └──────────┬─────────┘
│ │
└─────────────┬──────────────┘
▼
┌──────────────────────┐
│ Evaluation & │
│ Error Analysis │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ evaluation_summary │
│ .json │
└──────────────────────┘
MIT License. See LICENSE for details.
If you use this benchmark or methodology in research, please cite this repository.