🧩 Are Retrieval-Augmented Generation (RAG) Explanations Stable? Evaluating the Robustness of Rule-Based Reasoning under Perturbations
While Retrieval-Augmented Generation (RAG) models have advanced multi-hop question answering by grounding generation in retrieved evidence, the stability of their internal logical explanations under input perturbations remains poorly understood.
This repository provides an end-to-end framework to evaluate whether rule-based explanations generated by RAG systems remain stable when faced with:
- Document Deletion (subset context combinations across top-$k$ retrieved items).
- Context Reordering (permutations of retrieved documents).
- Question Paraphrasing (meaning-preserving semantic variations of the user prompt).
Using HotpotQA (distractor setting), dense vector retrieval (SentenceTransformers all-MiniLM-L6-v2 + FAISS IndexFlatIP), LLM generation (gpt-4o-mini), and strict LLM-as-a-Judge evaluation (gpt-4o), we evaluate 456 perturbation trials across 50 multi-hop reasoning baseline questions.
The pipeline operates across 8 sequential stages:
-
Data Selection (
src/prepare_data.py): Filters multi-hop factual questions from HotpotQA validation split with ground truth supporting facts. -
Indexing (
src/build_index.py): Extracts unique passage context and indexes document embeddings into a FAISS vector database using cosine similarity. -
Retrieval (
src/retrieve.py): Performs top-$k$ ($k=4$ ) dense vector retrieval for candidate queries. -
Baseline Answer Generation (
src/generate_answer.py): Generates concise baseline answers and verifies Exact Match /$F_1$ scores against gold answers. -
Base Rule Generation (
src/generate_explanations.py): Prompts the LLM to output structured JSON containing the answer and a logical rule/reasoning step. -
Perturbation Module (
src/run_perturbations.py): Executes 14 document deletion subsets, 3 context permutations, and 2 question paraphrases per baseline query. -
Stability Evaluation (
src/evaluate_stability.py): Evaluates explanation logic equivalence using a SOTA judge model (gpt-4o) and tracks prediction-level vs. explanation-level stability. -
Export & Visualization (
src/export_results.py&src/dashboard.py): Generates metric figures, Excel spreadsheets, HTML reports, and an interactive Streamlit GUI dashboard.
Across 456 perturbation trials, we track three primary metrics:
- Rule Stability (%): Percentage of trials where explanation logic remains fundamentally equivalent.
- Answer Stability (%): Percentage of trials where predicted answers match the baseline.
- Hidden Instability (%): Cases where the model gives the same answer but uses different/hallucinated logic.
| Category | Total Trials | Rule Stability (%) | Answer Stability (%) | Hidden Instability (%) |
|---|---|---|---|---|
| Overall | 456 | 73.03% | 74.12% | 7.24% |
| Document Deletion | 336 | 67.56% | 68.75% | 8.04% |
| Context Reordering | 72 | 87.50% | 90.28% | 6.94% |
| Question Paraphrasing | 48 | 89.58% | 87.50% | 2.08% |
- Answer accuracy doesn't guarantee reasoning faithfulness. Across 456 perturbation trials, I found that 7.24% of cases (and nearly 10% of cases where the model got the right answer) exhibited Hidden Instability — the model preserved its correct final answer while silently fabricating or breaking the logical rule behind it.
- Explanations are far more fragile than answers under context loss. When supporting documents were deleted, Rule Stability dropped to 67.56%, compared to 87–90% under reordering and paraphrasing — showing that removing information, not just rearranging or rephrasing it, is what actually breaks the model's reasoning.
- The model rarely admits uncertainty. When it did produce a wrong answer, 93.2% of the time it was a confident, specific, fabricated guess rather than an honest "I don't know" (only 6.8% of failures were graceful abstentions).
- Lexical and positional robustness is strong. The model's answers and rules were highly stable under document reordering (90.28% / 87.50%) and question paraphrasing (87.50% / 89.58%), suggesting reasoning is tied to semantic content rather than surface phrasing.
- Bottom line: correctness-only evaluation, the industry standard for RAG systems, would have silently missed every one of these failures, since the final answer looked right in each case.
.
├── README.md # Project documentation & execution guide
├── config.yaml # Model, retriever, and experiment configuration
├── requirements.txt # Complete Python dependency list
├── .env.example # Template for API keys (OpenAI / OpenRouter)
├── .gitignore # Files ignored by version control
├── run_pipeline.py # Master CLI pipeline runner
├── src/ # Modular Python source modules
│ ├── prepare_data.py # Stage 1: HotpotQA candidate filtering
│ ├── build_index.py # Stage 2: FAISS vector indexing
│ ├── retrieve.py # Stage 3: Dense vector retrieval
│ ├── generate_answer.py # Stage 4: Baseline RAG answer generation
│ ├── generate_explanations.py # Stage 5: Structured base rule extraction
│ ├── run_perturbations.py # Stage 6: Perturbation experiment execution
│ ├── evaluate_stability.py # Stage 7: LLM-as-a-Judge rule stability evaluation
│ ├── calc_metrics.py # CLI summary metric calculator
│ ├── export_results.py # Figure, Excel, and HTML report generator
│ └── dashboard.py # Interactive Streamlit GUI dashboard
├── data/ # Data directory
│ ├── raw/ # Raw dataset storage
│ └── processed/ # Pre-processed sample datasets & evaluation results
└── reports/ # Visual artifacts & documentation
├── figures/ # High-resolution chart exports (.png)
├── diagrams/ # Architecture & workflow diagrams (.png)
├── literature_review.md # Literature review documentation
└── final_report.html # Self-contained HTML report
- Python 3.8+
- Git
Clone the repository and install dependencies:
git clone https://github.com/omrusman/RAG-Explanation-Stability.git
cd RAG-Explanation-Stability
# Create a virtual environment (recommended)
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# Install required packages
pip install -r requirements.txtCopy the .env.example file to .env and set your API key:
cp .env.example .envEdit .env:
OPENAI_API_KEY=your_actual_openai_api_key_here(Optionally set OPENROUTER_API_KEY if using OpenRouter).
To launch the Streamlit visualization interface:
streamlit run src/dashboard.pyOpen your browser at http://localhost:8501.
The dashboard allows you to:
- Inspect overall Rule vs. Answer stability metrics across perturbation categories.
- Select any question to view the original text, ground truth, retrieved passages, and base explanation.
- Compare baseline rules side-by-side with perturbed rules (Deletion, Reordering, Paraphrase).
You can run the entire pipeline end-to-end using the master runner script:
python run_pipeline.py --step allOr execute individual pipeline stages separately:
# 1. Prepare candidates
python run_pipeline.py --step prepare
# 2. Build FAISS index
python run_pipeline.py --step index
# 3. Dense Retrieval
python run_pipeline.py --step retrieve
# 4. Generate Baseline Answers
python run_pipeline.py --step answer
# 5. Generate Explanation Rules
python run_pipeline.py --step explain
# 6. Run Perturbations
python run_pipeline.py --step perturb
# 7. Evaluate Stability
python run_pipeline.py --step evaluate
# 8. Export Reports & Figures
python run_pipeline.py --step reportThis README was structured and generated with the assistance of Gemini.



