A domain-specific medical Retrieval-Augmented Generation (RAG) system that answers health-related questions strictly based on retrieved evidence from a local vector database. Built with Python 3.12, PubMedBERT embeddings, ChromaDB, and a dual-LLM backend (Groq API + Local Qwen 2.5).
Core principle: If the answer is not in the database, the bot says so. No hallucination. No general-knowledge guessing.
- π Evidence-Based Answers β Retrieves from NHS, MedlinePlus, and Mayo Clinic data before generating any response.
- π‘οΈ Safety Agent β Automatically detects emergency symptoms (chest pain, severe bleeding, etc.) and returns an urgent-care disclaimer instead of a casual answer.
- π€ Telegram Bot Interface β Async bot using
python-telegram-botv20+ with non-blocking RAG processing. - π§ Dual LLM Backend β Switch instantly between Groq API (fast, high-quality) and local Qwen 2.5 0.5B (free, offline, private).
- π Local Vector Store β ChromaDB with PubMedBERT embeddings (
pritamdeka/S-PubMedBert-MS-MARCO) for medically-aware semantic search. - β‘ Lazy Initialization β Heavy models load only once and only when needed, preventing circular imports and double-loading.
βββββββββββββββ βββββββββββββββββββ ββββββββββββββββββββ
β Telegram ββββββΆβ main.py ββββββΆβ ClinicRAGPipelineβ
β User β β (Entry Point) β β (Orchestrator) β
βββββββββββββββ βββββββββββββββββββ ββββββββββ¬ββββββββββ
β
βββββββββββ¬βββββββββββ¬βββββββββββββ¬βββββββ
βΌ βΌ βΌ βΌ
ββββββββββββ ββββββββ ββββββββββββ ββββββββββββββ
β Safety β βRouterβ βRetriever β β Generator β
β Agent β β β β(ChromaDB)β β(Groq/Qwen) β
ββββββββββββ ββββββββ ββββββ¬ββββββ βββββββ¬βββββββ
β β
ββββββββββββΌββββββββββ β
β PubMedBERT β β
β Vector Embeddings β β
ββββββββββββββββββββββ β
β
ββββββββββββββββΌβββββββββββββββ
β NHS / MedlinePlus / Mayo β
β (JSON chunks in ChromaDB) β
βββββββββββββββββββββββββββββββ
| Component | Technology |
|---|---|
| Language | Python 3.12 |
| Embeddings | sentence-transformers + pritamdeka/S-PubMedBert-MS-MARCO (768-dim) |
| Vector DB | ChromaDB (chromadb>=0.5) with cosine HNSW indexing |
| Local LLM | Qwen 2.5 0.5B Instruct (Qwen/Qwen2.5-0.5B-Instruct) via HuggingFace Transformers |
| Cloud LLM | Groq API (llama-3.3-70b-versatile) |
| Telegram | python-telegram-bot>=20 (async) |
| Config | Custom YAML/JSON config loader + python-dotenv |
| Logging | Centralized clinic_rag_bot logger |
git clone <your-repo-url>
cd clinic_rag_botpython -m venv venv
# Windows
venv\Scripts\activate
# macOS/Linux
source venv/bin/activatepip install -r requirements.txtKey packages installed:
torch>=2.0.0
sentence-transformers>=3.0.0
transformers>=4.40.0
chromadb>=0.5.0
python-telegram-bot>=20.0
groq>=0.9.0
python-dotenv>=1.0.0
numpy>=1.23.5,<2.5.0
httpx>=0.27.0
Create a .env file in the project root:
# βββ Telegram βββ
TELEGRAM_BOT_TOKEN=your_telegram_bot_token_here
# βββ LLM Provider Toggle βββ
# Set to 'true' to use free local Qwen (offline, no API key needed)
# Set to 'false' to use Groq API (faster, higher quality)
USE_LOCAL_LLM=false
# βββ Groq (required if USE_LOCAL_LLM=false) βββ
GROQ_API_KEY=gsk_your_groq_key_here
# βββ HuggingFace (optional, for higher download limits) βββ
HF_TOKEN=your_huggingface_token
# βββ Chroma Collection Name βββ
# Must match between ingestion and retriever
CHROMA_COLLECTION=clinic_dataPlace your JSON files in:
data/raw/medlineplus/
data/raw/mayo/
data/raw/nhs/
Each JSON file should follow this schema:
{
"title": "Headaches",
"source": "NHS",
"url": "https://www.nhs.uk/conditions/headaches/",
"content": "Most headaches will go away on their own..."
}Run ingestion:
python src/ingestion.pyYou should see:
Loaded 5 documents from medlineplus
Ingested 47 chunks from 5 documents in 'medlineplus'
Total documents in collection: 120
# Using Groq (default)
python main.py --query "i have a headache"
# Using Local Qwen (free, offline)
python main.py --query "what are diabetes symptoms"# Correct way
python main.py --bot
# β Do NOT run the bot file directly (causes circular import)
# python bot/telegram_bot.pyThe bot will print:
Telegram Bot is starting...
Bot connected: @YourBotName
Bot is now polling. Send a message on Telegram!
clinic_rag_bot/
βββ .env # Environment variables
βββ .gitignore
βββ README.md # This file
βββ requirements.txt
βββ main.py # Entry point (CLI + Telegram)
β
βββ bot/
β βββ telegram_bot.py # Telegram adapter (async handlers)
β
βββ src/
β βββ pipeline.py # ClinicRAGPipeline orchestrator
β βββ pubmed.py # GeneralKnowledgeRetriever
β βββ retrievers.py # LocalRAGRetriever wrapper
β βββ vector_store.py # ChromaDB client (explicit embeddings)
β βββ embeddings.py # PubMedBERT wrapper
β βββ generator.py # AnswerGenerator (Groq / Qwen toggle)
β βββ local_llm.py # Qwen 2.5 0.5B local inference
β βββ ingestion.py # JSON β ChromaDB indexer
β βββ safety_agent.py # Emergency symptom detection
β βββ router.py # Query intent classification
β βββ config_loader.py # YAML/JSON config parser
β βββ logger.py # Centralized logging
β
βββ data/
βββ raw/
β βββ medlineplus/ # *.json medical articles
β βββ mayo/
β βββ nhs/
βββ chroma_db/ # Persistent ChromaDB (auto-generated)
ChromaDB defaults to all-MiniLM-L6-v2 (384-dim) for auto-embedding. We override this by computing embeddings with EmbeddingModelLocal (PubMedBERT, 768-dim) and passing them explicitly via collection.add(embeddings=...). This prevents the catastrophic "distance=31" mismatch bug.
ChromaDB returns distance, not similarity:
distance = 0β identicaldistance = 2β completely opposite- Conversion:
similarity = 1 - (distance / 2)
main.py does not instantiate ClinicRAGPipeline at import time. It uses a _get_pipeline() singleton. This prevents:
- Double model loading when
telegram_bot.pyimportsprocess_request - Circular import crashes
Groq's free tier has a 6,000β12,000 TPM limit. Medical documents are long, so the generator limits evidence to:
- Max 2 chunks
- Max 800 characters per chunk
- Max ~1,600 total context characters
This keeps prompts under ~1,500 tokens, avoiding 413 Request too large errors.
| Problem | Cause | Solution |
|---|---|---|
ImportError: cannot import name 'process_request' |
Circular import | Run python main.py --bot, never python bot/telegram_bot.py |
413 Request too large |
Prompt exceeds Groq TPM limit | Context is auto-truncated in generator.py. Lower top_k in retriever. |
401 Invalid API Key |
Groq key missing or wrong | Check .env file. Or set USE_LOCAL_LLM=true for free local mode. |
[] empty retriever results |
Wrong collection name | Ensure pubmed.py collection name matches ingestion (clinic_data) |
Distance values like 31.45 |
DB embedded with wrong model | Delete data/chroma_db and re-run python src/ingestion.py |
| Bot starts then exits immediately | run_bot() not called or loop not blocked |
Use asyncio.Event().wait() in telegram_bot.py |
This bot is not a substitute for professional medical advice, diagnosis, or treatment. It is an information retrieval tool designed to surface verified medical content from public databases.
- Always seek the advice of your physician or other qualified health provider with any questions you may have regarding a medical condition.
- If you are experiencing a medical emergency, call your local emergency services immediately.
This project is open-source. Feel free to fork and adapt for educational or research purposes.
Built with β€οΈ for responsible AI in healthcare.