AI-Powered Recommendation & Search Engine for Indian Standards (BIS)
Maanak is an intelligent Retrieval-Augmented Generation (RAG) platform designed to answer queries regarding Bureau of Indian Standards (BIS). Instead of relying on a language model's pre-trained knowledge—which can lead to hallucinations or outdated information—Maanak retrieves verified standard-specific data prior to generating precise, structured answers grounded in actual document content.
Maanak operates across three primary decoupled pipelines: Data Augmentation, Vector Retrieval, and Contextual Generation.
+-----------------------------------------------------------------------------------+
| 1. DATA AUGMENTATION PIPELINE |
| [ BIS Official Data ] -> [ Data Cleaning ] -> [ Structuring ] -> [ Embeddings ] |
+-----------------------------------------------------------------------------------+
|
v
+-----------------------------------------------------------------------------------+
| 2. RETRIEVAL PIPELINE |
| [ User Query ] -> [ Query Embedding ] -> [ Qdrant DB Search ] -> [ Top-5 Chunks] |
+-----------------------------------------------------------------------------------+
|
v
+-----------------------------------------------------------------------------------+
| 3. GENERATION PIPELINE |
| [ Top-5 Chunks + Query ] -> [ Prompt ] -> [ Grok LLM ] -> [ Structured Answer ] |
+-----------------------------------------------------------------------------------+
- Data Collection: Collects raw Indian Standards data directly from official sources.
- Data Cleaning & Structuring: Irrelevant noise is removed, and meaningful content is organized into a structured JSON format.
- Embedding Generation: Structured entries are processed through a
sentence-transformersmodel to generate high-dimensional numerical vector representations that capture semantic meaning.
- Vector Storage: Embeddings and metadata are stored in Qdrant Vector Database, supporting combined semantic similarity and keyword-based filtering.
- Query Embedding: Incoming user queries are embedded using the exact same sentence-transformer model to map into the shared vector space.
- Top-K Retrieval: Executes hybrid search in Qdrant to pull the top 5 (
k=5) most relevant context entries for the given query.
- Prompt Engineering: Merges the retrieved top-5 context entries with the original query into a structured system prompt.
- LLM Inference: Sends the context-rich prompt to the Grok (xAI) language model.
- Response Formatting: Post-processes the raw model output into a standardized layout for presentation.
- High Accuracy: Reduces hallucinations by grounding answer generation inside official BIS documentation.
- Scalable Knowledge Base: New standards can be ingested into Qdrant independently without retraining or fine-tuning models.
- Fast Retrieval: Qdrant vector indexing isolates only the relevant information entries instead of processing raw document trees.
- Decoupled Architecture: Modular design allows swapping generation models or vector databases without altering frontend or pipeline logic.
Maanak/
├── .github/
│ └── workflows/ # CI/CD automation and test pipelines
├── backend/ # FastAPI (Python) Service
│ ├── app/
│ │ ├── api/ # Route endpoints and request validation
│ │ ├── core/ # Configuration and environment management
│ │ ├── config/ # Configuration Files
│ │ ├── schema/ # Hold the Data transfer Object files
│ │ └── service/ # Transformer embeddings and LLM integrations
│ ├── tests/ # Unit and integration test suites
│ └── requirements.txt # Python dependencies
├── frontend/ # Next.js (TypeScript) Client
│ ├── public/ # Static assets, branding, and images
│ ├── src/
│ │ ├── components/ # UI elements (search, standards viewer, footer)
│ │ ├── pages/ # Application routes and dynamic views
│ │ ├── services/ # API integration client
│ │ └── styles/ # Global styles and layout themes
│ ├── package.json # Frontend scripts and dependencies
│ └── tsconfig.json # TypeScript configuration
├── extension/ # Chrome browser extension module
├── CONTRIBUTIONS.md # Contribution guidelines
└── README.md # Project documentation
- Python 3.9 or higher
- Node.js v18 or higher
- Running instance of Qdrant Vector Database (Local or Cloud)
- API Key for Groq
- Change directory into
backend:
cd backend- Create and activate a virtual environment:
uv init
source venv/bin/activate- Install required Python packages:
uv add -r requirements.txt- Configure environment variables (create a
.envfile inbackend/):
QDRANT_SERVER_URL=http://localhost:6333
QDRANT_API_KEY=qdrant_api_key_if_using_cloud_version
QDRANT_COLLECTION_NAME=bis_standards
GROQ_API_KEY=groq_api_key- Start the FastAPI backend server:
uv run python -m app.main- Open a new terminal and navigate to
frontend:
cd frontend- Install client dependencies:
npm install- Configure local environment variables (create
.env.localinfrontend/):
NEXT_PUBLIC_API_URL=http://localhost:8080/api- Run the development server:
npm run devAccess points:
- Web Client:
http://localhost:3000 - API Documentation:
http://localhost:8080/docs
Thank you for exploring Maanak. Contributions, feedback, and pull requests are welcomed to help advance automated access to Indian Standards.