A smart document-based study assistant built with Local Embeddings, Vector Search (FAISS), and Gemini AI. This app allows users to upload documents (PDFs, DOCX, PPTX, images), extract content using OCR, embed it locally, store it in a vector store, and ask natural language questions. Features a hybrid architecture combining cloud AI intelligence with local processing for privacy and speed.
A practical implementation of Retrieval-Augmented Generation (RAG) with a privacy-first approach.
Core Workflow
-
Upload Documents (PDF, DOCX, PPTX, images via OCR)
-
Extract & Process Text locally using various parsers and Tesseract OCR
-
Chunk & Embed using Local SentenceTransformers (all-MiniLM-L6-v2)
-
Store in FAISS vector database for semantic search
-
Receive Query → Search vector store for relevant context
-
Generate Answer using Gemini AI with retrieved context
-
Create Study Materials (MCQs, flashcards, summaries)
ai-study-assistant/
│
├── 📄 streamlit_app.py # Streamlit frontend interface
│
├── 📁 backend/ # FastAPI backend
│ ├── main.py # FastAPI application & routes
│ ├── config.py # Configuration settings
│ ├── faiss_store.py # FAISS vector database operations
│ ├── embeddings.py # Local embedding generation (SentenceTransformers)
│ ├── qa_engine.py # Q&A with Gemini AI & context retrieval
│ ├── mcq_generator.py # MCQ generation using Gemini
│ ├── flashcard_maker.py # Flashcard creation
│ ├── summarizer.py # Text summarization
│ ├── file_extractor.py # Document text extraction (PDF, DOCX, PPTX, OCR)
│ ├── rag_pipeline.py # RAG processing pipeline
│
│
├── 📁 vectorstore/ # FAISS vector store (auto-created)
│ └── faiss_index/ # Index and metadata storage
│
├── 📁 uploaded_docs/ # Uploaded documents storage (auto-created)
├── 📄 .env # Environment variables
├── 📄 .gitignore # Git ignore file
├──📄 README.md
└── requirements.txt
git clone https://github.com/m-sameerkhan/ai-study-assistant.git
/ai-study-assistant.git
cd ai-study-assistant
# Windows
python -m venv venv
venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
Tesseract OCR (for image text extraction)
-
Download from: https://github.com/UB-Mannheim/tesseract/wiki
-
Add to system PATH
Create a .env file in the project root:
# Gemini API Key (Get from: https://ai.google.dev/)
GEMINI_API_KEY=your_gemini_api_key_here
# Gemini Model (Use gemini-pro for free tier)
GEMINI_CHAT_MODEL=gemini-pro
# Local Embeddings Model
EMBEDDING_MODEL_NAME=all-MiniLM-L6-v2
# Storage Directories
FAISS_DIR=./vectorstore/faiss_index
UPLOAD_DIR=./uploaded_docs
# Text Processing
CHUNK_SIZE=800
CHUNK_OVERLAP=150
MAX_RESPONSE_TOKENS=500
TEMPERATURE=0.2
# Server Configuration
HOST=0.0.0.0
PORT=8000
DEBUG=True
API_BASE=http://127.0.0.1:8000
uvicorn backend.main:app --host 0.0.0.0 --port 8000 --reload
# Open new terminal
streamlit run streamlit_app.py
-
Upload PDF, DOCX, PPTX, TXT, or images (PNG, JPG)
-
PDFs/DOCX/PPTX: Extracted using specialized libraries
-
Images: Processed with Tesseract OCR
-
Text is cleaned, chunked, and prepared for embedding
-
Each text chunk is converted to embeddings using all-MiniLM-L6-v2
-
Embeddings are stored in FAISS vector database locally
-
Metadata preserves document references and chunk relationships
-
User question is embedded using the same local model
-
FAISS performs similarity search to find relevant document chunks
-
Top-k most relevant chunks are retrieved as context
-
Retrieved context + user question sent to Gemini AI
-
Gemini generates accurate, context-aware responses
-
Confidence scores indicate answer reliability
-
MCQs: Generate multiple-choice questions with explanations
-
Flashcards: Create Q/A pairs for active recall practice
-
Summaries: Produce concise summaries in various formats
- Upload your textbook PDFs, lecture notes, or research papers
- System extracts text and creates searchable embeddings
Q: "What are the key differences between supervised and unsupervised learning?"
A: [Based on your uploaded machine learning textbook]
Supervised learning uses labeled data with clear input-output pairs...
- MCQs: "Generate 5 questions about neural networks"
- Flashcards: "Create flashcards for Actuators chapter 2"
- Summaries: "Summarize this 10-page article in bullet points"
-
fastapi,uvicorn(Web framework & server) -
sentence-transformers(Local embedding models) -
faiss-cpu(Vector similarity search) -
google-generativeai(Gemini AI integration) -
pypdf2,python-docx,python-pptx(Document parsing) -
pytesseract,pillow,opencv-python(Image OCR processing) -
python-multipart(File upload handling) -
streamlit(Web interface) -
requests(API communication)
| Issue | Solution |
|---|---|
ModuleNotFoundError: No module named backend |
Run commands from the project root directory |
| Permission denied saving FAISS index | Run terminal as Administrator or fix folder permissions |
| Gemini API errors | Check API key validity and quota in Google AI Studio |
| Tesseract OCR not working | Ensure Tesseract is installed and added to PATH |
| Slow first run | SentenceTransformers downloads model (~90MB) on first use |
| Port already in use | Change port in .env or run with --port 8001 |
Muhammad Sameer Khan
LinkedIn
Contributions and forks are welcome!
If you'd like to extend this chatbot with:
- Additional document format support
- More export options (Anki, PDF, etc.)
- Mobile-responsive UI improvements
- Docker deployment scripts
- Performance optimizations
- Additional local LLM options
Feel free to fork or open a pull request.