Skip to content

Changes prakhar - #18

Merged
RK-NerdyBirdy merged 15 commits into
GDGVIT:devfrom
Prakhar-Sethi012:Changes_Prakhar
Aug 24, 2026
Merged

RK-NerdyBirdy merged 15 commits into
GDGVIT:devfrom
Prakhar-Sethi012:Changes_Prakhar

Conversation

@Prakhar-Sethi012

Copy link
Copy Markdown
Contributor

Overview

This PR finalizes the implementation of the jina-clip-v1 multimodal AI model and the new nltk query expansion logic. It resolves several critical infrastructure and database conflicts that occurred during the upgrade, ensuring the Docker containers build successfully and the Celery workers can safely process multimodal files.

Changes Implemented

1. Infrastructure & Docker Optimization

  • Resolved Rust Compilation Failures: Downgraded the Dockerfile from python:3.13 to python:3.11-slim. This allows pip to use pre-compiled binaries for HuggingFace tokenizers, entirely bypassing the maturin/cargo build crashes.
  • Fixed Startup Deadlocks: Removed the hardcoded Google DNS (8.8.8.8) in docker-compose.yml that was blocking GitHub connections.
  • Optimized Image Build: Baked the NLTK dictionaries (punkt, wordnet, omw-1.4) directly into the Docker image via the Dockerfile. The FastAPI server now boots instantly instead of hanging on runtime downloads.

2. Dependency Management

  • Added nltk>=3.9.1 to requirements.txt to support the new expansion.py logic.

3. Database Migration (pgvector)

  • Dimension Clash Resolved: Updated the SQLAlchemy FileContent model so the embedding column expects 768 dimensions instead of 384.
  • Schema Upgrade: Executed a manual SQL migration (ALTER TABLE ... TYPE vector(768)) and safely truncated legacy 384-dimension vectors to prevent Celery from crashing with sqlalchemy.exc.DataError.

Testing & Metrics

  • File Ingestion: Successfully tested end-to-end asynchronous indexing using Swagger UI.
  • Multimodal RAM Profiling: Monitored worker resources during image ingestion via docker stats.
    • The Celery worker safely peaked at ~2.0 GiB of RAM during the initial loading and inference phase of the Jina Vision/Text encoders.
    • Server remains well within the Docker memory limits.

This update completely changes how we handle file uploads by moving the heavy AI work into the background so the user doesn't have to wait.

Here is everything that changed:
- Added Redis and Celery worker to docker-compose to handle background jobs.
- Updated Dockerfile to cache Python packages for faster builds.
- Fixed the database migrations and added a 'state' column to track file progress (processing, indexed, failed).
- Fixed the circular import crashes in the database models.
- Updated the upload route to instantly send jobs to the queue and return a task ID.
- Added a new '/status' live stream endpoint to show real-time progress to the frontend.
- Added an automatic cleanup step in the worker so temporary files are always deleted.
- Brought back the GET /files/ endpoint so the dashboard can list all uploaded documents.
…l security/data bugs

Infrastructure & Architecture:
- Completely decoupled file uploading from the main API thread using Redis and Celery.
- AI (sentence-transformers) now runs safely in background workers.
- Implemented Server-Sent Events (SSE) for live frontend status tracking.
- Restored unified hybrid search endpoint.

Security & Vulnerability Patches:
- Patched 'Zip Slip' directory traversal vulnerability in archive extraction.
- Secured upload route against path traversal and malicious overwrites using UUID prefixing.
- Prevented temp-directory memory leaks on corrupted ZIP uploads.

Data Integrity & Database:
- Resolved race conditions in Celery workers using IntegrityError try/rollback blocks.
- Forced strict failure propagation so Celery correctly reports 'FAILURE' instead of swallowing exceptions.
- Added 'state' column to File model to track indexing progress.
- Added 'relation_type' to FileRelationship model.
- Fixed one-directional relationship bug by implementing symmetric ORM queries.
- Cleaned up app/database namespace collisions.
…d migrate vector DB

Infrastructure & Build Fixes:
- Downgraded Dockerfile to Python 3.11-slim to bypass Rust/Cargo compilation errors for HF tokenizers.
- Removed hardcoded DNS from docker-compose.yml to resolve network timeouts when pulling NLTK dictionaries.
- Baked NLTK data (punkt, wordnet, omw-1.4) directly into the Docker image to eliminate startup hangs.

AI & Database Migration:
- Added missing nltk dependency for expansion.py.
- Updated FileContent ORM model to Vector(768) to support the new Jina CLIP embedding model.
- Executed database migration to stretch the pgvector column and cascade-delete legacy 384-dimension files.
@RK-NerdyBirdy
RK-NerdyBirdy merged commit ed88170 into GDGVIT:dev Aug 24, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants