A robust, advanced Video Transcriber system that extracts the exact voice from video files or YouTube URLs using Whisper transcription.
- Upload or link lecture videos (YouTube supported)
- Automatic audio extraction via FFmpeg/yt-dlp
- Speech-to-text transcription using Whisper
- Streamlit web UI for an interactive experience
- CLI for terminal usage
- Strict adherence to original voice (no LLM alteration/hallucination)
graph TD
A[Video Input File / YouTube URL] -->|yt-dlp / FFmpeg| B(Audio Extraction)
B --> C[16kHz Mono WAV Audio File]
C -->|faster-whisper + WhisperX| D(Speech Recognition & Alignment)
D --> E[Exact Voice Transcript JSON]
E --> F[Display in Streamlit UI / CLI Output]
- Python 3.10+: Core language
- Streamlit: Web interface
- faster-whisper: High-performance speech-to-text
- WhisperX: Forced alignment and punctuation
- FFmpeg & yt-dlp: Media processing and downloading
- Video Upload/Link: User provides a video file or YouTube URL.
- Audio Extraction: Audio is extracted and converted to a standardized WAV format (16kHz, mono).
- Transcription & Alignment:
faster-whispertranscribes the audio to text, andWhisperXhandles forced alignment to capture the exact spoken voice with accurate timestamps. - Display: The final, unadulterated transcript is displayed directly to the user.
- Clone the repository:
git clone https://github.com/hammad986/video-transcriber.git cd video-transcriber - Set up Python environment:
python -m venv my_env source my_env/Scripts/activate # On Windows: my_env\Scripts\activate.bat pip install -r requirements.txt
- Install System Dependencies:
- FFmpeg: You must have
ffmpeginstalled and added to your system PATH.
- FFmpeg: You must have
- Start the Streamlit app:
streamlit run transcriber/app.py
- Open your browser:
- Go to http://localhost:8501
- Upload a video or enter a YouTube URL.
- Run the CLI tool:
python main.py
- Follow the on-screen prompts to input your video path.
- Transcription:
faster-whisper(small model by default, configurable invideo_qa/config.py). - Alignment:
whisperxalignment models. - Important Note: No local LLM (like Ollama) or cloud LLM models are used in this version. The system is purely focused on extracting the exact voice from the video without AI corrections, alterations, or summarization.