DeepDoc is a full-stack Retrieval-Augmented Generation (RAG) platform that enables users to upload PDF documents and engage in context-grounded, interactive Q&A sessions with AI. Built on Next.js 15, Azure OpenAI, Pinecone Vector DB, and Drizzle ORM with Neon PostgreSQL, DeepDoc extracts, semantically chunks, vectorizes, and analyzes complex PDF files with extreme accuracy.
- 📑 Smart PDF Upload & Storage: Secure file upload with Vercel Blob storage, client-side validation (4MB size limits, PDF MIME check), and server-side text extraction.
- 🧠 Dynamic Semantic Chunking: Advanced NLP pipeline that splits documents by semantic sentence shifts using embedding distance quantiles rather than arbitrary character splits.
- ⚡ Vector Search & Namespace Isolation: Vector embeddings (
text-embedding-3-small) stored in Pinecone vector indexes, isolated by unique per-document namespaces. - 🛡️ Grounded AI Answers: Powered by Azure OpenAI (
gpt-4o-mini) with strict anti-hallucination system prompting (answers strictly based on document context). - 🔄 Smart Deduplication & Token Packing: Query context ranking with similarity score thresholds, substring deduplication, and token budget packing.
- 🖥️ 3-Pane Interactive Workspace:
- Left: Sidebar with past PDF chats and quick navigation.
- Center: Embedded live PDF viewer powered by Google Docs Viewer.
- Right: Real-time chat interface with optimistic updates via TanStack Query.
flowchart TD
User([User]) -->|Upload PDF| UploadComp[PDFUpload Component]
UploadComp -->|Server Action| PDFProc[lib/pdf-process.ts]
subgraph Storage & Database
PDFProc -->|Store PDF File| VercelBlob[(Vercel Blob Storage)]
PDFProc -->|Save Chat Metadata| NeonDB[(Neon PostgreSQL via Drizzle)]
end
subgraph Semantic Chunking Pipeline
PDFProc -->|Extract Text| PDFParse[pdf-parse]
PDFParse -->|Split Sentences| NaturalNLP[natural SentenceTokenizer]
NaturalNLP -->|Sliding Windows| EmbedGen[text-embedding-3-small]
EmbedGen -->|Cosine Distance| ShiftDetect[Quantile Shift Detection]
ShiftDetect -->|Merge Chunks| FinalChunks[Semantic Chunks]
end
FinalChunks -->|Upsert Vectors + Metadata| Pinecone[(Pinecone Vector DB)]
subgraph Query & RAG Flow
User -->|Ask Question| ChatRoute[app/api/chat/route.ts]
ChatRoute -->|Embed Query| QueryEmbed[text-embedding-3-small]
QueryEmbed -->|Similarity Query| Pinecone
Pinecone -->|Top Matching Vectors| ContextRet[lib/context.ts]
ContextRet -->|Dedupe & Budget Pack| GroundedContext[Context Block]
GroundedContext -->|Prompt + Context| AzureModel[gpt-4o-mini]
AzureModel -->|Grounded Response| User
ChatRoute -->|Persist History| NeonDB
end
Unlike naive chunkers that split text every
-
Sentence Tokenization: Cleans whitespace and splits raw extracted text into sentences using
natural'sSentenceTokenizer, preserving common English honorifics (Mr.,Dr.,etc.). -
Context Windowing: Groups sentences with adjacent neighbors (default
bufferSize = 2) to maintain context. -
Vector Distance Shift Detection: Generates embeddings for each sentence buffer using Google's
text-embedding-004and computes the cosine distance ($1 - \text{cosineSimilarity}$ ) between sequential sentence vectors usingmathjs. -
Quantile Thresholding: Identifies topic boundaries by calculating quantile cutoffs (e.g., 90th percentile distance threshold via
d3-array). -
Group & Merge: Groups sentences into chunks at boundary points and selectively merges smaller adjacent chunks if their combined length is under limits and semantic similarity is high (
$> 0.9$ ).
- Extracted chunks are converted into 768-dimensional embeddings.
- Vector IDs are generated deterministically using MD5 hashes (
md5(chunk.text)). - Vectors and metadata (
text,startIndex,endIndex,chatId,fileKey) are upserted into Pinecone in batches under a dedicated namespace (fileKey) for strict document isolation.
When a question is submitted:
- Embedding: The query is converted into a vector using
text-embedding-004. - Top-K Vector Query: Pinecone queries the document namespace for top matching vectors (default
topK = 20). - Score Filtering & Deduplication: Matches with similarity scores below the threshold (default
0.7) are discarded.dedupeRankedMatchesremoves redundant chunks whose normalized text is a substring of a higher-ranked match. - Token Budget Packing:
packIntoTokenBudgetconcatenates the most relevant unique chunks until reaching a strict token budget (default750 tokens/~3000 chars), truncating gracefully if needed.
- The retrieved context block is wrapped into a strict system prompt for
gemini-2.5-flash. - Anti-Hallucination Constraints: The model is instructed to answer strictly using the provided context block. If the context is insufficient or off-topic, it is constrained to respond with a standardized message.
-
Resiliency:
callGeminiWithRetry()features exponential backoff retry logic and a 20-second timeout handling transient upstream API issues ($503, 500, 429, 408$ ). -
History Persistence: Both user questions and AI system responses are persisted to PostgreSQL
messagestable via Drizzle ORM.
PostgreSQL database schema managed via Drizzle ORM:
| Field | Type | Description |
|---|---|---|
id |
serial (PK) |
Unique auto-incrementing chat session ID |
pdfName |
text |
Original filename of the uploaded PDF |
pdfUrl |
text |
Public URL of the stored PDF on Vercel Blob |
fileKey |
text |
Unique storage key / Pinecone namespace |
createdAt |
timestamp |
Timestamp when the chat was created |
| Field | Type | Description |
|---|---|---|
id |
serial (PK) |
Unique message ID |
chatId |
integer (FK) |
Reference to chats.id |
content |
text |
Content of the message |
role |
pgEnum('system', 'user') |
Role of the message sender |
createdAt |
timestamp |
Timestamp when message was created |
- Framework: Next.js 15 (App Router), React 19, TypeScript
- Database & ORM: Neon Serverless PostgreSQL (
@neondatabase/serverless), Drizzle ORM (drizzle-orm,drizzle-kit) - Vector Database: Pinecone (
@pinecone-database/pinecone) - AI Models & Embeddings: Azure OpenAI (
openai)- LLM:
gpt-4o-mini - Embeddings:
text-embedding-3-small
- LLM:
- File Storage: Vercel Blob (
@vercel/blob) - NLP & Mathematics:
pdf-parse,natural(SentenceTokenizer),d3-array(quantiles),mathjs(vector operations),md5 - Client State & Data Fetching: TanStack React Query (
@tanstack/react-query), Axios - Styling & UI: Tailwind CSS, Lucide React, Radix UI Primitives,
react-hot-toast
Create a .env file in the root directory:
# Database
DATABASE_URL="postgresql://user:password@neon-db-host/dbname?sslmode=require"
# Azure OpenAI Configuration
AZURE_OPENAI_API_KEY="your_azure_openai_api_key"
AZURE_OPENAI_ENDPOINT="https://your-resource-name.openai.azure.com/"
AZURE_OPENAI_API_VERSION="2024-06-01"
AZURE_OPENAI_CHAT_DEPLOYMENT_NAME="gpt-4o-mini"
AZURE_OPENAI_EMBEDDING_DEPLOYMENT_NAME="text-embedding-3-small"
# Pinecone Vector DB
PINECONE_API_KEY="your_pinecone_api_key"
PINECONE_INDEX="chatpdf"
# Vercel Blob Storage Token
BLOB_READ_WRITE_TOKEN="your_vercel_blob_token"
# Optional RAG Tuning
CONTEXT_TOP_K=20
CONTEXT_SCORE_THRESHOLD=0.7
CONTEXT_MAX_TOKENS=750npm install
# or
pnpm installnpx drizzle-kit pushnpm run devOpen http://localhost:3000 in your browser.
npm run dev: Starts the Next.js development server.npm run build: Builds the production bundle.npm run start: Runs the built production server.npm run lint: Runs Next.js ESLint checks.