SignAI is a real-time sign language recognition and translation system for German Sign Language (DGS). The production path is a BiLSTM+attention seq2seq model over MediaPipe Holistic keypoints (sentence-level gloss translation); in parallel, signai/word_classification/ is developing a 3-stream fusion model (pose heatmap CNN + hand DINOv3 transformer + mouth DINOv3 transformer) for single-word classification — see Single-Word Classifier. The project won 1st place at the Jugend forscht state competition and received coverage in SZ, BR, and other media outlets.
Primary languages: Python (core, app), CSS/HTML/JavaScript (product website).
- SignAI — Sign Language Translator
SignAI turns a short clip of someone signing into text. The desktop app is the usual front door, but the same request-level flow applies wherever a video reaches the inference API:
- Capture —
app/camera.pyrecords webcam video in the desktop app. Pressing Record starts capture; pressing it again stops and hands the clip off for translation. - Upload —
app/api_call.pysends the video to the local Flask API (POST /api/uploadonhttp://127.0.0.1:5000). - Preprocess —
api/signai_api.pysaves the upload todata/live/video/, thenapi/preprocessing_live_data.pyruns MediaPipe Holistic over every frame and writes the extracted keypoints todata/live/live_dataset.csv. - Infer —
api/inference.pyloads the trained.kerasmodel (see Environment Variables for how the model path is chosen) together with the gloss tokenizer, and predicts a translation with a confidence score. - Display — the API returns JSON to the app, which renders the result in the main window.
Webcam ──▶ camera.py ──▶ api_call.py ──▶ signai_api.py ──▶ preprocessing_live_data.py
│
▼
app UI ◀── inference.py + gloss_tokenizer.json
The model itself — a BiLSTM encoder feeding an LSTM decoder with multi-head attention — is trained separately ahead of time; see Models & Training and Architecture for how that training happens and what the network looks like.
- OS: Windows (primary target). macOS/Linux support is in development and not yet verified end-to-end.
- Python: no version is pinned in
requirements.txt, but the desktop build tooling (app/builds/README.md) documents and is tested against Python 3.10–3.12; that's the range to use unless you're prepared to debug version issues yourself. - Hardware: a webcam for live recognition. A GPU is recommended for training and speeds up inference; CPU-only works but is slower.
- Disk: at least 5 GB free — trained models, caches, and MediaPipe assets add up quickly.
git clone https://github.com/Stefanos0710/SignAI.git
cd SignAI
python -m venv venv
venv\Scripts\activate # macOS/Linux: source venv/bin/activate
pip install -r requirements.txt
requirements.txt pins the versions that are actually load-bearing here —
notably tensorflow==2.16.2, keras==3.7.0, mediapipe==0.10.21,
protobuf==4.25.8, and numpy==1.26.4. These aren't arbitrary: newer
protobuf or numpy releases break MediaPipe or TensorFlow compatibility, so
avoid upgrading them individually.
Every command below assumes the virtual environment from the previous step
is active. The cwd column matters — several scripts resolve paths (like
data/train_data) relative to the current working directory, not to their
own location.
| Component | Command | Run from | Port |
|---|---|---|---|
| Desktop app | python app.py |
app/ |
— (talks to the API internally) |
| Flask inference API | python -m api.signai_api |
repo root | 5000 |
| Web API (Flask + SocketIO) | python main.py |
repo root | 8000 (override with PORT) |
| Product website | python main.py |
product_webside/ |
5000, bound to 0.0.0.0 |
| Letter classification demo site | python app.py |
signai/letter_classification/website/ |
5000 |
Note that the Flask inference API and both demo websites default to the same port (5000) — don't try to run more than one of them at a time without changing the port in code, or you'll get a bind conflict.
The desktop app doesn't launch the API as a separate process you start
yourself; it imports and drives api/signai_api.py directly through
app/api_call.py. Running python -m api.signai_api by hand is mainly
useful for testing the API in isolation (e.g. with curl or Postman)
outside the desktop UI.
| Variable | Effect | Default |
|---|---|---|
SIGNAI_MODEL_PATH / SIGNAI_MODEL |
Overrides which .keras model the API loads |
newest models/trained_model_v*.keras |
SIGNAI_DISABLE_SITE_CLEANUP |
Set to 1 to stop the app/API from stripping user site-packages off sys.path at startup |
cleanup enabled |
PORT |
Port for the main.py Flask-SocketIO server |
8000 |
The site-packages cleanup exists because a system-wide protobuf install can
silently shadow the pinned protobuf==4.25.8 from the venv and break
MediaPipe — disable it only if you're sure your environment doesn't have
that conflict. Desktop packaging has its own separate set of build-only
environment variables, documented in app/builds/README.md.
Primary translation model — BiLSTM encoder + LSTM decoder with 8-head MultiHeadAttention.
python signai/sentence_classification/train.py
Configuration is at the bottom of signai/sentence_classification/train.py (defaults: version 38.4, 200 epochs, batch 64, multi_attention).
Key features:
- Mixed precision training (
mixed_float16global policy) - Per-epoch WER, BLEU-1..4, ROUGE-1/2/L evaluation
- Epoch-wise augmentation (temporal: stretch/warp/freeze/dropout; spatial: shift/scale/rotate/noise), implemented in
signai/sentence_classification/augmentation.py - Transformer architecture also available in
signai/sentence_classification/experimental_transformer.py
Latest trained models:
| Version | Type | Notes |
|---|---|---|
| v36 | BiLSTM-Seq2Seq | Latest internal version — June 2026 |
| v30 | Seq2Seq | Latest public version — April 2026 |
| v29 | Seq2Seq | 200+ epochs, full history |
| v28 | Seq2Seq | 200+ epochs, full history |
- Vocabulary: 800+ gloss tokens
- Output length: Up to 15 tokens per sentence
- Input features: 426 per frame (7 pose + 21 left hand + 21 right hand + 93 face landmarks, each x/y/z)
Three-stream fusion model, signai/word_classification/. Own dataset/preprocessing/training — independent of the sentence pipeline.
| Step | Script | Output |
|---|---|---|
| 1. Download | download.py |
Public DGS Corpus (eaf + openpose) |
| 2. Segment | segmentation_videos.py (needs ffmpeg) |
per-word clips, dataset/word_clips/ |
| 3. Preprocess | preprocessing.py |
dataset/processed/{train,val,test}_{data,images}.npz |
| 4. Cache hand/face features | models/hand_stream.py / models/face_stream.py --extract-features {split} |
*_hand_features.npz / *_face_features.npz |
| 5. Train | models/train.py |
checkpoints/word_classifier_best.keras |
All run from the repo root. Step 4 needs torch + transformers and a Hugging Face token with the DINOv3 ViT-S/16 license accepted (huggingface-cli login or HF_TOKEN) — steps 3 and 5 don't.
preprocessing.py: MediaPipe pose/hand keypoints → shoulder-center + shoulder-scale → Savitzky–Golay smooth → per-landmark Gaussian heatmap (x/y only, z dropped) + left-hand/right-hand/mouth RGB crops (Shades-of-Gray color correction) → resample/pad to 32 frames.
Model — models/model.py::WordClassifier:
| Stream | Input | Encoder | Output |
|---|---|---|---|
| Pose | (32, 49, 96, 96) heatmaps | SlowOnly-R50 3D CNN (ResNet-50, inflated) → Dense | 384-d |
| Hands | 2×(8, 384) frozen DINOv3 ViT-S/16 (cached) | shared pre-LN transformer (2 blocks, 6 heads, CLS) | 768-d |
| Mouth | (8, 384) frozen DINOv3 ViT-S/16 (cached) | own transformer head, same design | 384-d |
Pose(384) + Hands(768) + Mouth(384) = concat(1536) → Dense(512, ReLU) → Dropout(0.3) → Dense(classes)
DINOv3 runs once offline into the cache; training only touches cached features + Keras layers, no torch needed at train time. models/fusion.py = same architecture, computes pose heatmaps on the fly instead of from cache.
Dataset: DGS Corpus word clips. In progress — no published accuracy yet.
An independent fingerspelling-alphabet classifier under signai/letter_classification/ (previously the standalone SignAlphaSet sub-project). Has its own dataset, models, and a small Flask demo site — does not share code or training data with the sentence/word classifiers above.
python signai/letter_classification/train.py
Dataset download and preprocessing: signai/letter_classification/download.py, preprocess_v2.py/preprocess_v3.py. All scripts in this subtree must be run from the repo root — their data/model paths are hardcoded relative to it (e.g. signai/letter_classification/data/...).
Sentence classification's training CSVs live in data/train_data/ (parsed cache .parsed_cache.pkl, delete or pass --rebuild-cache to re-parse; CSVs git-ignored, only example_for_train_data.csv tracked). The word classifier is separate — see its own dataset layout in Single-Word Classifier.
MediaPipe (Holistic and/or Face Mesh) is used for keypoint extraction, but each consumer has its own pipeline — none of these three are interchangeable.
| Script | Purpose | Output | Landmarks |
|---|---|---|---|
signai/preprocessing/train_data.py |
Sentence classification training data | 426 features/frame (×3 xyz) | 7 pose + 42 hand + 93 face |
api/preprocessing_live_data.py |
Live inference | 151 features (frame-averaged) | 543 landmarks × 2 (xy) |
signai/word_classification/preprocessing.py |
Word classifier training data (3-stream) | per-landmark Gaussian heatmaps (32,49,96,96) + hand/mouth RGB crops (32,224,224,3) | 7 pose + 21+21 hand |
Normalization pipeline:
- Video-wise shoulder midpoint centering
- Shoulder-distance scaling
- Savitzky–Golay temporal smoothing (window 9, polyorder 2)
- Linear interpolation for missing keypoints
Encoder: Input(426) → Dense(1024) → LayerNorm → Dropout → DepthwiseConv1D → BiLSTM(512) → LayerNorm
Decoder: Embedding(256) → LayerNorm → LSTM(512) → LayerNorm → MultiHeadAttention(8 heads, residual) → Concat → Dense(512) → Dropout → LayerNorm → Dense(vocab, softmax)
Pose heatmaps (32,49,96,96) → SlowOnly-R50 3D CNN → Dense(384) ─┐
Hand features 2×(8,384) → shared transformer head → concat(768) ─┼─ concat(1536) → Dense(512, ReLU) → Dropout(0.3) → Dense(classes)
Mouth features (8,384) → transformer head → (384) ─────────┘
- Workflow: Press Record → perform signs → press again → upload to API → display translation
- Result display:
QPlainTextEdit, hidden until ready, shows translation with optional debug info - Single-instance lock: TCP port 52391
- Logging: stdout/stderr tee'd to
logs/desktop_app.log - Settings:
app/settings/settings.json - Path handling:
resource_path()for bundled assets,writable_path()for per-user data (%LOCALAPPDATA%\SignAI\) - Qt fix:
fix_qt_plugin_path()must run before any PySide6 import - User-site cleanup: User site-packages stripped from
sys.pathto avoid protobuf version conflicts - Build: PyInstaller spec at
app/SignAI - Desktop.spec, output atbuild/SignAI - Desktop/SignAI - Desktop.exe
- PyInstaller spec:
app/SignAI - Desktop.spec— bundles models, tokenizers, UI, icons (pathex set to repo root) - Updater:
app/start_updater.py, spec atapp/SignAI - Updater.spec - Build scripts:
app/builds/build-exe.py(--onefile,--include-models,--clean,--dry-run),build-updater-exe.py,build-final-app.py,build-zip.py— seeapp/builds/README.mdfor the full release sequence and its own build-only environment variables - Runtime API overrides: see Environment Variables
- Camera feed: If no image appears, press "Switch Camera" repeatedly. Close other camera-using apps.
- Admin privileges: Some operations may require elevation. Future releases will reduce this.
- First-run delay: Models load from disk on first launch — wait a few seconds for the UI to become responsive.
- Recognition quality: Degrades for casual or atypical signing. Addressed by planned augmentation and larger datasets.
- Improve accuracy 3x via full datasets, larger compute, synthetic augmentation, and transformer architectures
- Expand vocabulary to thousands of gloss tokens
- Reduce admin access requirements
- Natural language rendering (gloss → grammatical sentences)
- Multilingual support (ASL planned)
- Fork and create a branch:
git checkout -b feat/my-change - Add tests and documentation for changes
- Open a Pull Request with a clear description
- Do not commit large model binaries — use release assets
Non-commercial license. See LICENSE. Contact maintainers for alternative arrangements.
- General / press: hello@signai.dev
- Support: open an issue at GitHub Issues