Target Developer/Creative: Joshua Maina (mainajoshua713@gmail.com)
System Classification: B2B Live Event Accessibility & Cognitive Augmentation Appliance
Engineering Architecture: Hexagonal / Clean Architecture (Hard Real-Time Streaming)
VerbaClear is an ambient, AI-powered real-time vocabulary assistant designed as a B2B event technology solution. The platform targets a deep psychological pain point: the public anxiety, cognitive overload, and loss of focus experienced by audience members when speakers use intense, advanced vocabulary or niche industry jargon.
Unlike traditional translation frameworks that output heavy, scrolling lines of raw text, VerbaClear works as an educational intelligence filter within the same language (English-to-Simple-English). It processes live speech invisibly, extracting only uncommon vocabulary and serving instantaneous, low-distraction synonyms to the entire room, preserving audience engagement and speaker flow.
The system is fully documented following professional Computer Science and enterprise software engineering standards:
| Document | Purpose & Scope | Link |
|---|---|---|
| System Architecture (SAD) | Signal processing theory, Bloom filter math, C4 models, ADRs, latency budget, and STRIDE security analysis. | Docs/ARCHITECTURE.md |
| Software Requirements (SRS) | Formal Functional (FR-1 to FR-6) and Non-Functional Requirements (NFR-1 to NFR-4) with acceptance test matrices. | Docs/SPECIFICATION.md |
| Data Dictionary & Protocols | SQLite DDL schemas (WAL mode), JSON Schema contracts, and TypeScript wire protocols for WebSockets. | Docs/DATA_DICTIONARY.md |
| Phased Execution Blueprint | End-to-end processing pipeline flowcharts and strict sequential build verification gates. | Docs/SYSTEM_SPECIFICATION.md |
| Original Spec PDF | Archived reference documentation from project inception. | Docs/VerbaClear_Project_Documentation.pdf |
[Live Presenter Audio] (16kHz PCM)
│
▼
[Silero VAD v5] ── (Noise / Silence Gating) ──► Dropped (<1ms)
│ (Speech Detected, 250ms trailing pause)
▼
[faster-whisper Engine] (CTranslate2, distil-large-v3) ──► Latency: ~150ms
│ (Transcribed Text + Word Timestamps)
▼
[Lexical Brain] (NGSL Bloom Filter + spaCy Lemmatizer) ──► Latency: ~5ms
│ (Uncommon Terms Filtered: Rank > 3,500)
▼
[SQLite Lexicon Store] (Curated Synonyms + Phonetics) ──► Latency: ~2ms
│
▼
[FastAPI WebSocket Hub] ──► Dispatch: ~3ms
├──► /ws/stage ──► Stage Lower-Third Overlay (7s Spring Decay Card)
└──► /ws/audience ──► Mobile Companion PWA (Definitions + Anki Export)
TOTAL END-TO-END LATENCY: ~460ms - 600ms (P99 < 800ms)
To optimize comprehension without causing visual distraction, VerbaClear segments data presentation into two specific tiers:
| Feature Component | Primary Stage Presentation Screen | Audience QR Companion Portal |
|---|---|---|
| Target Intent | Subtle, split-second ambient glance. | Deep-dive contextual look and bookmarking. |
| Format Layout | Isolated lower-third overlay card featuring 1-3 punchy synonyms (e.g., Labyrinthine → Highly Complex). Fades out in 7 seconds. | Full mobile web environment displaying synonym + 1-sentence dictionary definition + phonetics. |
| Interaction | Zero interaction. Completely passive. | Active. Attendees tap to save words to a personal spaced-repetition deck (e.g., Anki, CSV, flashcards). |
Development progresses consecutively across 6 discrete phases with hard exit gates:
- Phase 1: The Auditory Ear (Audio & Speech Pipeline): 16kHz PortAudio ring buffer + Silero-VAD v5 +
faster-whisperCTranslate2.
Exit Gate: Live mic speech transcribes to console in < 800ms with zero silence hallucinations. - Phase 2: The Semantic Brain (Lexical Intelligence): NGSL/NAWL in-memory Bloom filter + spaCy lemmatizer + SQLite curated lexicon.
Exit Gate: Complex words in live speech output punchy synonyms, ignoring top 3,500 words. - Phase 3: The Nervous System (Event Hub & Gateway): FastAPI async server hosting background audio thread +
/ws/stage&/ws/audienceWebSockets.
Exit Gate: Real-time speech triggers synchronized JSON payloads to connected WebSocket clients within 50ms. - Phase 4: The Ambient Eye (Stage Lower-Third Display): Next.js / React 19 + Framer Motion chroma-key overlay with 7s decay.
Exit Gate: Spoken rare words appear on second monitor / OBS lower-third and auto-fade cleanly. - Phase 5: The Attendee Notebook (Audience QR Companion): Mobile PWA with live feed, IndexedDB local persistence, and Anki (
.apkg) / CSV export.
Exit Gate: QR scan on smartphone shows live feed; bookmarking exports importable Anki deck. - Phase 6: Production Hardening (AV & Venue Release): Physical mixer audio testing, 100% offline network verification, and Docker CLI orchestrator.
Exit Gate: 30-minute keynote test with 50 connected mobile devices, zero dropped frames, < 1.2s latency.
VerbaClear/
├── Docs/ # Comprehensive engineering specifications
│ ├── ARCHITECTURE.md # System design, C4 diagrams, ADRs, latency budget
│ ├── SPECIFICATION.md # SRS with functional & non-functional requirements
│ ├── DATA_DICTIONARY.md # SQLite DDL, WebSocket schemas & TypeScript types
│ ├── SYSTEM_SPECIFICATION.md # Technical pipeline & sequential execution plan
│ └── VerbaClear_Project_Documentation.pdf
├── src/
│ ├── domain/ # Pure business models & abstract interfaces (Ports)
│ │ ├── models.py # AudioChunk, TranscribedSegment, StageOverlayCard
│ │ └── interfaces.py # AudioSourcePort, SpeechToTextPort, LexicalFilterPort
│ ├── infrastructure/ # Framework-specific adapters
│ │ ├── audio/ # PortAudio driver & Silero-VAD v5 wrapper
│ │ ├── asr/ # faster-whisper CTranslate2 worker
│ │ ├── nlp/ # NGSL Bloom filter & spaCy lemmatizer
│ │ ├── storage/ # SQLite WAL lexicon repository
│ │ └── export/ # Anki .apkg, WebVTT, and SRT transcript exporters
│ ├── application/ # Core orchestrator pipeline & metrics
│ ├── api/ # FastAPI ASGI application & WebSockets
│ └── web/ # Self-contained offline presentation interfaces
│ ├── admin/ # Broadcast AV control room console
│ ├── stage/ # OBS / vMix transparent lower-third overlay
│ └── companion/ # Attendee mobile PWA with offline IndexedDB & Service Worker
├── scripts/ # Appliance automation & launchers
│ ├── install_appliance.sh # Hardware provisioning & model verification
│ └── run_appliance.sh # Single-command AV booth runtime launcher
├── systemd/ # 24/7 background service configuration
│ └── verbaclear.service # Production systemd unit file
├── tests/ # Comprehensive test suite (unit, integration, E2E)
├── pyproject.toml # Modern Python project & dependency configuration
└── README.md
Provision a dedicated Linux AV machine (Ubuntu/Debian) with system PortAudio libraries and models:
./scripts/install_appliance.shStart the real-time processing engine and broadcast hub:
./scripts/run_appliance.shOnce running, the appliance exposes four zero-install surfaces on the venue local network:
- AV Operator Control Room (
/admin): Hardware input device selection, real-time signal oscilloscope, Silero VAD gating, persistent stage blackout toggle, manual card injection, domain context pack upload/management, CEFR vocabulary demographic slider (B1/B2/C1/C2), NDI stream controls, and post-event transcript downloads. - Stage Lower-Third Overlay (
/stage): Transparent chroma-key window for OBS Studio, vMix, and secondary stage displays. Auto-fades cards with 7-second spring animations. - Native NDI & HTTP Alpha Stream (
/api/broadcast/stream/alpha&/api/broadcast/frame): Zero-halo 1080p RGBA transparent video broadcast output for NewTek NDI video switchers and direct OBS/vMix inputs. - Attendee Mobile Companion PWA (
/companion): Accessible via QR code (/api/session/qr). Operates 100% offline via Service Worker & IndexedDB, supporting real-time card streaming, search, bookmarks, and one-tap Anki (.apkg) export. - Post-Event Analytics & Subtitle Transcripts (
/api/export/transcript&/api/export/session-report): Instant export to WebVTT (.vtt), SubRip (.srt), plain text (.txt), and full readability & simplification analytics JSON for video editors.
VerbaClear includes a fast, zero-dependency command line interface:
# Start appliance with live microphone audio and specific context pack
verbaclear --start-audio --pack fintech --port 8000
# Inspect and enumerate all host audio hardware devices
verbaclear devices
# Run appliance diagnostic check (models, PortAudio drivers, database, directories)
verbaclear diagnose