A FastAPI-based service for real-time speech-to-text using faster-whisper.
stt_preview.mov
# Install system requirements
sudo apt install libcublas-12-9 libcudnn9-cuda-12
# Install python dependencies
python3 src/setup.py
source src/stt-venv/bin/activateStart the service:
cd src/
python app.pyPython example:
import requests
with open("audio.wav", "rb") as f:
response = requests.post(
"http://localhost:47102/transcribe",
files={"file": f}
)
print(response.json()["text"])| Method | Path | Description |
|---|---|---|
| GET | /health |
Check service health and loaded model |
| POST | /transcribe |
Transcribe audio, with optional segments, word timestamps, or translation |