diff --git a/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/final-report.md b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/final-report.md new file mode 100644 index 0000000..c0ef470 --- /dev/null +++ b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/final-report.md @@ -0,0 +1,67 @@ +# Final Report: LearningAI architecture books admin brain and revision cleanup + +## Outcome + +The architecture library, learner Admin view, fullscreen tutor flow, and +learning-book generation contract now describe and expose the same product +model. The copy is clearer, implementation status is explicit, and the UI uses +existing data rather than invented analytics. + +## Accepted Results + +- Rewrote the Tutor System Architecture and User Brain Architecture books with + clearer chapter structure, linked sources, expanded terminology, and honest + completion boundaries. +- Reframed local beta controls as operator controls and explained why each + retained control exists. +- Made the Admin learner view the default, reduced the primary navigation to + four decision-useful sections, and added per-learner concept and evidence + inspection. +- Added fullscreen tutor chat with or without a PDF, including responsive + mobile behavior. +- Changed generated learning-book guidance and fallback material to produce + explanations, relationships, worked examples, diagrams, questions, and code + when the subject requires it. +- Prevented stale chapter audio from attaching to a rewritten chapter title. + +## Rejected Results + +- Did not claim that aspirational brain contracts are fully enforced by the + runtime. +- Did not fabricate secure tenant identity or server-wide learner analytics; + the current Admin selector truthfully uses locally stored learner names. +- Did not delete ambiguous untracked directories or unrelated dirty-worktree + edits. + +## Conflicts Resolved + +- Preserved the existing deep Admin debugging routes under a collapsed + advanced section so operational tests and specialist workflows remain + available without dominating the main dashboard. +- Kept the established Dexie and learning-book schemas unchanged to avoid an + unnecessary data migration. + +## Verification Evidence + +- Full test program: 275 Node plus 591 DOM tests, 866 total, passed. +- Type checks, production build, formatting, whitespace checks, and the + Graphify scratch guard passed through `brain:postchange`. +- Desktop and 390x844 mobile Chrome checks passed for the revised library, + status chapter, Admin learner view, design component map, operator controls, + and no-PDF fullscreen chat. +- The local application returned HTTP 200 at `http://127.0.0.1:3001/`. + +## Remaining Risks + +- Per-learner Admin data is local-device data keyed by a display name, not + authenticated multi-tenant storage. Production use still requires durable + user IDs, authorization, row-level isolation, consent, and audit controls. +- Rewritten chapters intentionally ignore old title-mismatched audio. New audio + assets should be generated before considering every chapter audio-ready. +- Some runtime brain behavior remains less strict than the documented target; + the books label those areas as partial or deferred. + +## Reusable Follow-up + +- Use this report and `state.json` as the acceptance record for future + multi-user persistence, audio regeneration, and evidence-gating work. diff --git a/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/orchestration.md b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/orchestration.md new file mode 100644 index 0000000..4b723f8 --- /dev/null +++ b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/orchestration.md @@ -0,0 +1,47 @@ +# Orchestration: LearningAI architecture books admin brain and revision cleanup + +## Execution Rules + +- Keep the original objective intact. +- Ask for approval before risky, expensive, external, or destructive actions. +- Keep immediate blocking work local. +- Delegate only bounded, disjoint, materially useful packets. +- Integrate packet results before final verification. + +## Branching Rules + +- If book copy overstates runtime completion, rewrite it as accepted direction + plus explicit implementation status. +- If a requested Admin metric has no real data source, do not fabricate it; + show an unavailable/needs-instrumentation state. +- If a cleanup candidate is referenced, persisted, or user-owned, retain it. +- If browser access remains blocked, use the existing rendered tests and a + local Playwright fallback and record the limitation. + +## Packet Prompts + +### Packet 1: Books + +Rewrite built-in architecture content for clear technical reading, linked +citations, glossary coverage, and truthful completion status. + +### Packet 2: Admin + +Reduce Admin to essential operational and learner-analysis views. Add readable +guidance and per-user brain inspection using real stored data. + +### Packet 3: Chat and Learning Books + +Enable no-PDF fullscreen chat and improve generated learning-book material with +explanations, examples, diagrams, and code when appropriate. + +### Packet 4: Cleanup and QA + +Audit changed and directly connected files, remove only proven dead code, and +run focused plus global verification. + +## Completion Audit + +- Every success criterion has source or test evidence. +- No unrelated dirty-worktree changes were reverted. +- Remaining runtime gaps are reported plainly. diff --git a/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/plan.md b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/plan.md new file mode 100644 index 0000000..c35647a --- /dev/null +++ b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/plan.md @@ -0,0 +1,103 @@ +# LearningAI architecture books admin brain and revision cleanup + +## Goal + +Turn the Revision library, architecture books, Admin dashboard, fullscreen chat, +and generated learning books into one coherent product that a motivated +teenager can understand without reducing the technical depth. + +## Success Criteria + +- Tutor System Architecture chapters are reorganized and rewritten in clear + language, with an explicit current-state/completion chapter. +- User Brain Architecture chapters are renamed and rewritten around the actual + learner model, evidence boundaries, goals, and implementation status. +- App Design Language component maps have readable spacing, linked citations, + a complete glossary, and a clear explanation or removal of local beta + controls. +- Admin shows only decision-useful views and explains what each view means. +- Admin can select a learner and inspect that learner's knowledge graph, + mastery, misconceptions, and learning patterns. +- Chat can become fullscreen even when no PDF is loaded. +- Generated learning books read like revision material with explanations, + examples, diagrams, and code where the studied subject needs them. +- Cleanup removes only verified dead/obsolete code or files and preserves all + unrelated user changes. +- Focused tests, typecheck, build, postchange, and rendered UI checks pass. + +## Current Context + +- The worktree is already dirty from voice architecture work and other user + changes; unrelated edits must be preserved. +- Graphify MCP is attached to the wrong project. Local + `graphify-out/graph.json` is the architecture authority. +- The connected in-app browser blocks both `localhost:3001` and + `127.0.0.1:3001` with `ERR_BLOCKED_BY_CLIENT`; rendered validation must use + the repo's Playwright/test fallback unless browser access recovers. +- `brain:postchange` exists and passes. The older `brain:debug` and + `brain:ui-regression` scripts referenced by the generic debug skill are not + present in this checkout. + +## Constraints + +- Follow AGENTS.md: Graphify first, then directly connected files only. +- Do not regenerate or manually edit `graphify-out`. +- Keep architecture graph concepts separate from the learner-facing brain. +- Keep Admin operational and scannable rather than turning it into a marketing + page. +- Do not delete files without evidence that they are unused or generated. + +## Risks + +- Architecture books may describe contracts the runtime does not yet enforce. +- AdminView and RevisionView are large shared surfaces with broad tests. +- Learning-book schema and Dexie migrations are high-risk boundaries. +- Fullscreen chat must not break PDF study mode or mobile layout. +- Broad cleanup could accidentally remove user-owned worktree changes. + +## Approval Required + +- No external writes or deployment. +- Destructive file deletion is gated by an evidence-backed cleanup report. No + ambiguous files will be deleted automatically. + +## Work Packets + +1. Architecture books and citations: + `src/lib/tutorBook.json`, `src/lib/userBrainArchitectureBook.ts`, + `src/views/RevisionView.tsx`, architecture readiness tests. +2. Design language and component map: + `src/views/RevisionView.tsx` and directly connected design components. +3. Admin simplification and per-user brain: + `src/views/AdminView.tsx`, learner-memory read APIs, navigation tests. +4. Fullscreen chat: + `src/views/StudyView.tsx`, `src/components/ChatPanel.tsx`, store/types and + rendered study-flow tests. +5. Generated revision books: + `src/memory/memory.orchestrator.ts`, learning-book types, Revision rendering, + and focused memory tests. +6. Cleanup and verification: + changed-work audit, dead-code evidence, tests, build, postchange, and + Playwright fallback. + +## Integration Policy + +Prefer existing schemas, state, components, and visual language. Keep edits +inside established ownership boundaries. Integrate only changes that are +supported by live source and focused tests. + +## Verification + +- Focused source-contract and rendered tests per packet. +- `npm run test:node` +- relevant Vitest DOM tests +- `npm run lint` +- `npm run build` +- `npm run brain:postchange -- --reason architecture-admin-revision-cleanup` +- rendered desktop/mobile Playwright checks with console and overflow evidence. + +## Reusable Artifacts + +- This workflow directory. +- A concise final report with accepted changes, cleanup evidence, and remaining + implementation gaps. diff --git a/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/results.md b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/results.md new file mode 100644 index 0000000..8e5a853 --- /dev/null +++ b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/results.md @@ -0,0 +1,33 @@ +# Results + +## Books + +- Tutor architecture: 13 rewritten chapters with current-state and release + chapters. +- User brain architecture: 8 rewritten chapters with a substantial glossary + and linked references. +- App design language: clearer component map and operator-control rationale. + +## Product Surfaces + +- Admin defaults to learner analysis and keeps specialist diagnostics under an + advanced disclosure. +- Tutor chat supports fullscreen use without a document. +- Learning books use revision-material structure rather than transcript-like + concept lists. + +## Cleanup + +- Removed proven `.DS_Store` artifacts and the generated `.tmp-test` + directory. +- Preserved unrelated dirty files, untracked `docs/`, `output/`, and the + separate voice workflow because ownership or obsolescence was not proven. + +## Verification + +- `npm test`: 866 passed. +- `npm run lint`: passed. +- `npm run build`: passed. +- `npm run brain:postchange -- --reason architecture-admin-revision-cleanup`: + passed. +- Desktop and mobile rendered checks: passed. diff --git a/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/state.json b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/state.json new file mode 100644 index 0000000..b67e0a8 --- /dev/null +++ b/.workflow/learningai-architecture-books-admin-brain-and-revision-cleanup/state.json @@ -0,0 +1,31 @@ +{ + "title": "LearningAI architecture books admin brain and revision cleanup", + "slug": "learningai-architecture-books-admin-brain-and-revision-cleanup", + "created_at": "2026-06-08T06:15:58+00:00", + "status": "completed", + "approval": { + "required": true, + "granted": true, + "notes": "User requested broad cleanup. Ambiguous destructive deletions remain gated by evidence." + }, + "packets": [ + { "id": "books", "status": "completed" }, + { "id": "admin", "status": "completed" }, + { "id": "chat-learning-books", "status": "completed" }, + { "id": "cleanup-qa", "status": "completed" } + ], + "verification": { + "status": "passed", + "checks": [ + "Local Graphify queries identified the connected architecture surfaces.", + "Focused architecture, audio, Admin, Study, and Revision tests passed.", + "Full npm test passed: 275 Node tests and 591 DOM tests, 866 total.", + "npm run lint passed.", + "npm run build passed.", + "npm run brain:postchange -- --reason architecture-admin-revision-cleanup passed.", + "git diff --check passed.", + "Desktop and mobile Chrome rendering checks passed for the library, architecture chapters, learner Admin, component map, operator controls, and fullscreen no-PDF chat.", + "The local server returned HTTP 200 at http://127.0.0.1:3001/." + ] + } +} diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/final-report.md b/.workflow/sub-300ms-voice-brain-architecture-research/final-report.md new file mode 100644 index 0000000..55525f7 --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/final-report.md @@ -0,0 +1,63 @@ +# Final Report: Sub 300ms Voice Brain Architecture Research + +## Outcome +Use a Tutor-owned realtime voice broker. Do not keep relying on Deepgram Voice Agent as the whole agent runtime, and do not treat MisoTTS as duplex. The new design should be: + +Browser AudioWorklet -> Tutor broker WebSocket -> Deepgram STT -> GPT-4o-mini learner stream -> MisoTTS streaming/cancellable audio -> browser playback + +In parallel: + +Tutor broker -> GPT-5.5 background agent/tool queue -> web/search/PDF/code/tool results -> foreground insertion queue + +## Accepted Results +- Thinking Machines' article supports the architecture split the user wants: a real-time interaction model stays present while a background model handles deeper reasoning/tool use and streams results back into the conversation. +- Their technical direction uses 200ms micro-turns and persistent streaming sessions. We should copy the interaction pattern, not pretend a cascaded provider stack has the same native architecture. +- Miso Labs claims 110ms realtime latency and local/on-prem deployment, but the public MisoTTS card identifies it as an 8B text-to-speech model based on Sesame CSM. It is the voice renderer, not the whole interaction model. +- Deepgram's own latency docs place streaming transcription at 150-300ms and total transcript latency at 200-500ms. This makes full user-speech-to-semantic-answer under 300ms unrealistic as a guaranteed system target with cloud STT plus cloud LLM. +- OpenRouter supports `openai/gpt-4o-mini` through an OpenAI-compatible API, making it suitable for the small foreground learner LLM. +- OpenAI's GPT-5.5 docs show streaming, function calling, structured outputs, and Responses API tools such as web search, file search, code interpreter, hosted shell, apply patch, MCP, and tool search. That fits the async background layer. +- Current LearningAI already has useful primitives: `/api/voice-agent`, browser microphone/audio playback, barge-in stopping, voice tools, `/api/tts`, and `scripts/misotts_api_server.py`. + +## Rejected Results +- Rejected "Miso solves duplex." MisoTTS is TTS-only in the public model card; duplex must be implemented in the broker. +- Rejected "all output under 300ms" as a hard guarantee. The source-backed target should separate local cancellation/reaction latency from full semantic answer latency. +- Rejected a big-bang replacement. Keep the current Deepgram Voice Agent path as fallback until the new broker is proven. + +## Conflicts Resolved +- The user's desired Thinking Machines-style full-duplex feel conflicts with Miso's half-duplex TTS nature. Resolution: use Miso only for audio rendering and keep STT open during playback; cancel/flush TTS on barge-in. +- The current code delegates the whole voice loop to Deepgram Voice Agent. Resolution: add a Tutor-owned `/api/voice-broker` beside it, then migrate after measured proof. + +## Verification Evidence +- Graphify query routed the current voice architecture to `server.ts`, `src/components/ChatPanel.tsx`, `src/lib/voiceAgentTools.ts`, `src/lib/chatAgentTools.ts`, `src/memory/brain.context.ts`, `src/memory/brain.rehearsal.ts`, `scripts/misotts_api_server.py`, `src/memory/beta.diagnostics.ts`, and `src/views/AdminView.tsx`. +- Targeted source inspection confirmed `server.ts` proxies `/api/voice-agent` to Deepgram Voice Agent, configures Deepgram listen, GPT-4o-mini think, and Deepgram speak providers, and has `/api/tts` routing to MisoTTS. +- Targeted source inspection confirmed `ChatPanel.tsx` streams PCM16, plays binary PCM, stops active audio on barge-in, and handles voice tool callbacks. +- Targeted source inspection confirmed `scripts/misotts_api_server.py` currently generates full WAV responses under a lock. + +## Remaining Risks +- Actual Miso whole-utterance output is verified on the Vast.ai machine, but it + is not yet a realtime streaming path. Cold load was 87.77s; warm whole-request + tunnel latency was 6.13s for a very short utterance. +- OpenRouter GPT-4o-mini TTFT and GPT-5.5 background latency must be measured in the chosen region. +- True Miso streaming may require modifying the Miso generator internals; a phrase-chunked bridge is the fallback. +- The current browser microphone path uses ScriptProcessor, which should be replaced by AudioWorklet for stable low-latency frames. +- Provider keys, GPU deployment, and live traffic are still pending user input. + +## Reusable Follow-up +Implementation order: +1. Done locally: add `VITE_VOICE_BROKER_MODE=custom` and `/api/voice-broker` beside the current path. +2. Done locally: stage brain-context, previous-memory, active-book/document metadata, foreground model, MisoTTS, and GPT-5.5 background queue events without provider traffic. +3. Next after GPU/keys: add AudioWorklet 20ms capture for the new broker. +4. Next after keys: implement STT-only Deepgram WebSocket client with interim/final transcript events. +5. Next after keys: stream GPT-4o-mini from OpenRouter with short voice prompts and no foreground blocking tools. +6. Next after GPU: upgrade Miso service to stream/cancel audio; begin with phrase chunking because the verified wrapper is whole-utterance WAV output. +7. Next after GPT-5.5 key: replace the staged background queue with real GPT-5.5 tool execution and result insertion. +8. Add Admin diagnostics for p50/p95/p99 stage timings. + +Minimal Vast.ai machine: +- 1x RTX 4090 24GB preferred, or RTX 3090 24GB if cheaper. +- 8 vCPU minimum, 12-16 preferred. +- 32GB RAM minimum, 64GB preferred. +- 100GB NVMe minimum. +- Ubuntu + CUDA 12.x + PyTorch-compatible image. + +Safer but not necessary for first beta: L40S 48GB or A100 40GB. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/orchestration.md b/.workflow/sub-300ms-voice-brain-architecture-research/orchestration.md new file mode 100644 index 0000000..b502609 --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/orchestration.md @@ -0,0 +1,26 @@ +# Orchestration: Sub 300ms Voice Brain Architecture Research + +## Execution Rules + +- Keep the original objective intact. +- Ask for approval before risky, expensive, external, or destructive actions. +- Keep immediate blocking work local. +- Delegate only bounded, disjoint, materially useful packets. +- Integrate packet results before final verification. + +## Branching Rules +- If Miso is not full-duplex, implement duplex at the broker layer. +- If sub-300ms full answer latency conflicts with STT/LLM/TTS vendor ranges, report the realistic latency target instead of overpromising. +- If implementation requires live keys or GPU access, stop at readiness and ask for the machine/key inputs. + +## Packet Prompts +- P1 external research: verify Thinking Machines, MisoTTS, Deepgram, OpenRouter, and GPT-5.5 facts. +- P2 current architecture: use Graphify first, then inspect only voice/socket/tool files. +- P3 target architecture: define broker protocol, cancellation, async background jobs, and latency budget. +- P4 implementation readiness: define GPU spec, config inputs, staged code plan, and verification gates. + +## Completion Audit +- Graphify-first source discovery completed. +- Packet notes completed. +- Final report completed. +- Runtime source changes intentionally deferred until GPU/API-key inputs are available. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/packets/P1-external-research.md b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P1-external-research.md new file mode 100644 index 0000000..d457543 --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P1-external-research.md @@ -0,0 +1,28 @@ +# P1 External Research + +## Objective +Research Thinking Machines' interaction-model architecture and the provider constraints for the proposed Tutor voice stack. + +## Sources +- Thinking Machines Lab, "Interaction Models: A Scalable Approach to Human-AI Collaboration", May 11, 2026. +- Miso Labs product page and MisoLabs/MisoTTS Hugging Face model card. +- Deepgram streaming STT latency and live audio documentation. +- OpenRouter GPT-4o-mini and server-side web search docs. +- OpenAI GPT-5.5 model and release docs. + +## Findings +- Thinking Machines' core split is exactly the product direction we want: a real-time interaction model stays present while an asynchronous background model performs longer reasoning, browsing, and tool calls, then streams results back into the live conversation when appropriate. +- Their native model uses 200ms micro-turns and persistent streaming sessions in GPU memory. We cannot reproduce that with a cascaded STT -> text LLM -> TTS stack, but we can mimic the behavior with a persistent broker, short audio frames, speculative acknowledgement, cancellable TTS, and async background jobs. +- MisoTTS is an 8B text-to-speech model based on Sesame CSM. It is not a full-duplex interaction model; it should be treated as the audio renderer for foreground text, not the conversational brain. +- Miso's 110ms claim is a vendor claim and should be measured locally after deployment. Current LearningAI wrapper returns a full WAV, so local code must be upgraded before we can test first-audio latency. +- Deepgram docs describe streaming latency as a combination of network, model processing, buffer size, and client processing. Their stated typical ranges put transcription latency at 150-300ms and total transcript latency at 200-500ms, with Flux EOT at 100-500ms. +- OpenRouter supports `openai/gpt-4o-mini` through an OpenAI-compatible API. This is a reasonable foreground learner LLM if we stream tokens and keep prompts short. +- GPT-5.5 is suitable for the background agent/tool layer, especially for web search, file/code analysis, PDF generation, and long-running reasoning. It should not be in the foreground first-audio path. + +## Decision +Build a two-layer broker: +- Foreground voice interaction loop: browser audio frames -> Tutor voice broker -> Deepgram STT partial/final -> GPT-4o-mini streaming -> MisoTTS streaming/cancellable audio -> browser. +- Background agent loop: intent detector/tool router -> GPT-5.5 or OpenRouter tool-enabled model -> web/search/PDF/code tools -> result packets -> foreground loop insertion queue. + +## Verification Notes +All provider claims must be benchmarked on the final GPU/region with p50/p95/p99 timings. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/packets/P2-current-architecture.md b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P2-current-architecture.md new file mode 100644 index 0000000..8394e4f --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P2-current-architecture.md @@ -0,0 +1,28 @@ +# P2 Current LearningAI Architecture + +## Objective +Map the current voice architecture through Graphify before source inspection. + +## Graphify Context +Graphify query found the relevant cluster: +- server.ts +- src/components/ChatPanel.tsx +- src/lib/voiceAgentTools.ts +- src/lib/chatAgentTools.ts +- src/memory/brain.context.ts +- src/memory/brain.rehearsal.ts +- scripts/misotts_api_server.py +- src/memory/beta.diagnostics.ts +- src/views/AdminView.tsx + +## Source Findings +- `server.ts` attaches `/api/voice-agent` as a WebSocket and currently proxies to `wss://agent.deepgram.com/v1/agent/converse`. +- The current Deepgram Voice Agent config sets listen provider to Deepgram Flux, think provider to `gpt-4o-mini`, and speak provider to Deepgram Aura. +- `ChatPanel.tsx` opens microphone capture, converts browser audio to PCM16, buffers frames until voice auth is accepted, sends frames over `/api/voice-agent`, plays returned PCM, and handles barge-in by stopping active audio buffer nodes. +- `ChatPanel.tsx` also handles `FunctionCallRequest` by executing client/local tools and sending `FunctionCallResponse` back through the socket. +- `src/lib/voiceAgentTools.ts` already defines the voice tool surface we need for the background layer: study context, graph updates, flashcards, answer evaluation, current page vision, diagram rendering, and web search. +- `scripts/misotts_api_server.py` is already a FastAPI wrapper for MisoTTS 8B but uses a lock and returns a generated WAV after `gen.generate(...)`, so it is not yet live-streaming or multi-session friendly. +- `/api/tts` can route to `miso-tts-8b` with `x-miso-tts-api-url`, but this is suitable for read-aloud/chat TTS, not true realtime voice yet. + +## Decision +Do not bolt Miso into the current Deepgram Voice Agent config. Replace the provider-owned `/api/voice-agent` loop with a Tutor-owned broker while preserving the browser voice UI, local tool functions, and diagnostics/event ledger shape. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/packets/P3-target-architecture.md b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P3-target-architecture.md new file mode 100644 index 0000000..dd30c2d --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P3-target-architecture.md @@ -0,0 +1,51 @@ +# P3 Target Architecture + +## Objective +Define the brokered, low-latency architecture. + +## Components +- Browser voice client: AudioWorklet microphone capture, 20ms PCM16 frames, jitter buffer for playback, barge-in/cancel messages, visual voice state. +- Tutor Voice Broker: FastAPI or Node WebSocket service responsible for session state, cancellation, latency metrics, STT stream, learner LLM stream, TTS stream, and background job queue. +- Deepgram STT client: persistent WebSocket using interim results, VAD events, and tuned endpoint/utterance settings. +- Learner LLM: OpenRouter `openai/gpt-4o-mini`, streaming, short prompt, no long tools in the first pass. +- MisoTTS service: local GPU service with warmed model, streaming/cancellable synthesis, phrase-level chunks, and audio-frame output. +- Background agent: GPT-5.5 tool runner for web search, code/file analysis, PDF creation, and long reasoning, writing results into a session event queue. + +## Message Shape +- client.audio_frame: binary PCM16 frame plus sequence/timestamp metadata where needed. +- client.cancel_audio: sent on barge-in. +- broker.transcript_partial / broker.transcript_final: Deepgram STT events normalized for UI and LLM. +- broker.assistant_text_delta: foreground GPT-4o-mini deltas. +- broker.audio_chunk: Miso PCM/Opus chunks. +- broker.background_job_started / broker.background_result_delta / broker.background_result_final. +- broker.metric: stage timings for audio capture, STT partial/final, LLM first token, TTS first audio, playback start, cancellation. + +## Duplex Strategy +Miso is half-duplex, so simulate duplex at the broker layer: +- Always keep STT open, even while TTS is playing. +- On user speech above barge-in threshold, send client.cancel_audio, abort current GPT stream if appropriate, cancel Miso generation, flush playback queue, and resume listening. +- Allow foreground model to issue short acknowledgements while background work runs. +- Never wait for background GPT-5.5 before first spoken answer. + +## Example Flow +User says: "Tell me about Pikachu, and also look up Nvidia prices." +- STT partial detects topic and current-data request. +- Foreground GPT-4o-mini starts: "Sure. Pikachu is an Electric-type Pokemon..." +- Broker starts background job: `web_search` or stock-price lookup for Nvidia. +- Miso streams the foreground explanation. +- Background result arrives with cited/latest Nvidia price. +- Foreground insertion policy waits for a natural clause boundary, then says: "Quick update on Nvidia: ..." + +## Latency Budget +- Browser frame: 20ms target. +- Browser -> broker: 20-80ms depending geography. +- Deepgram interim STT: aim for first useful partial under 150-300ms, but vendor docs show total transcript can be 200-500ms. +- GPT-4o-mini TTFT through OpenRouter: must be measured; target under 150ms p50 with warmed connection and short prompts, but do not budget it as guaranteed. +- Miso first audio: vendor target 110ms; must be measured after replacing full-WAV generation with streaming. +- Playback jitter buffer: 40-80ms. + +## Honest Target +Sub-300ms full semantic answer is not reliable with cascaded cloud STT + cloud LLM + local TTS. The realistic target is: +- p50 barge-in/cancel under 100ms after speech energy detection. +- p50 first acknowledgement audio around 300-600ms after end/commit, depending STT and LLM TTFT. +- background tool results inserted asynchronously without blocking the foreground tutor. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/packets/P4-implementation-readiness.md b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P4-implementation-readiness.md new file mode 100644 index 0000000..0956cd7 --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/packets/P4-implementation-readiness.md @@ -0,0 +1,40 @@ +# P4 Implementation Readiness + +## Objective +Prepare the first implementation slices and GPU recommendation. + +## Minimal GPU +Recommended starting Vast.ai machine: +- GPU: RTX 4090 24GB or RTX 3090 24GB. +- CPU: 8 vCPU minimum, 12-16 preferred. +- RAM: 32GB minimum, 64GB preferred. +- Disk: 100GB NVMe minimum. Miso weights are about 32.8GB F32 on Hugging Face, plus repo/env/cache/logs. +- Network: choose the lowest-latency region to the user and provider edges; test Tokyo/US-West depending where the app is used. +- CUDA: modern NVIDIA driver with CUDA 12.x compatible PyTorch. + +## Safer GPU +If rental cost is still acceptable, L40S 48GB or A100 40GB gives more room for FP16, KV cache, concurrency, and avoiding offload. This is not required for the first single-user test. + +## First Code Slices +1. Add a new broker path behind a feature flag, leaving existing Deepgram Voice Agent path intact. +2. Replace ScriptProcessor capture with AudioWorklet 20ms frames for the new broker path. +3. Implement broker STT-only Deepgram stream and normalize transcript events. +4. Add foreground GPT-4o-mini streaming turn manager with short acknowledgement policy. +5. Upgrade Miso service to stream chunks and support cancellation. If streaming requires deeper Miso generator changes, start with phrase-chunked HTTP plus immediate cancellation while measuring. +6. Move voice web/tool calls from client-only callback shape into server-side/background job queue where possible, preserving UI visual-focus events. +7. Add GPT-5.5 background job adapter with provider/model config and strict result insertion protocol. +8. Add diagnostics: p50/p95/p99 per stage, audio underruns, cancellation latency, provider errors. + +## Required Inputs From User +- Vast.ai instance URL/SSH details once rented. +- Deepgram API key. +- OpenRouter API key for GPT-4o-mini. +- GPT-5.5 API key/provider choice. +- Whether Miso should run on the same Node app host or a separate GPU FastAPI host. + +## Verification +- Unit tests for broker message normalization and tool-job routing. +- Local mock-provider test for the full socket protocol. +- Live single-user latency drill with timestamped stage metrics. +- Browser desktop/mobile smoke test of voice UI states and barge-in. +- `npm run lint` and `npm run build` after source changes. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/plan.md b/.workflow/sub-300ms-voice-brain-architecture-research/plan.md new file mode 100644 index 0000000..3da7727 --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/plan.md @@ -0,0 +1,60 @@ +# Sub 300ms Voice Brain Architecture Research + +## Goal +Design an implementation-ready voice brain architecture for Tutor/LearningAI that mimics the Thinking Machines interaction-model split: a low-latency foreground learner dialogue loop plus an asynchronous background reasoning/tool layer. + +## Success Criteria +- Source-backed understanding of Thinking Machines' interaction/background split. +- Source-backed constraints for Deepgram STT, MisoTTS, OpenRouter GPT-4o-mini, and GPT-5.5. +- Graphify-first map of the current LearningAI voice architecture and the narrow files likely to change. +- Proposed socket/message architecture that can support barge-in, cancellation, streaming STT, streaming text, streaming TTS, and background tool result insertion. +- Latency budget with honest feasibility notes for the user's sub-300ms target. +- Minimal Vast.ai GPU recommendation for MisoTTS with a safer headroom option. +- Clear implementation slices to start once the user provides the GPU endpoint and API keys. + +## Current Context +- Repo: /Users/mfuad16/Documents/LearningAI. +- Required navigation: Graphify first, then targeted source reads. +- Existing Graphify query routed voice work through server.ts, src/components/ChatPanel.tsx, src/lib/voiceAgentTools.ts, src/lib/chatAgentTools.ts, src/memory/brain.context.ts, src/memory/brain.rehearsal.ts, and scripts/misotts_api_server.py. +- Current implementation proxies browser PCM to Deepgram Voice Agent at /api/voice-agent. Deepgram currently owns STT, the GPT-4o-mini "think" provider, and Deepgram TTS. +- Current code already has /api/tts support for a local MisoTTS HTTP server and a scripts/misotts_api_server.py wrapper, but it returns complete WAV responses rather than a true incremental audio stream. +- Current voice tools already include look_at_study_context, update_graph, generate_flashcards, evaluate_answer, look_at_current_page, render_diagram, and web_search. + +## Constraints +- MisoTTS is TTS-only and half-duplex; duplex behavior must be implemented by our broker, not assumed from the TTS model. +- Deepgram STT latency alone can be 150-300ms processing plus network, so "full semantic answer under 300ms" is not realistic for every utterance; the achievable target is first reaction/backchannel under 300ms and first meaningful audio chunk as low as possible. +- Browser-to-cloud geography matters. Place the broker, Miso GPU, and external API egress in the closest feasible region to the user and Deepgram/OpenRouter edge. +- Avoid sending OpenRouter GPT-5.5 into the foreground voice loop. It belongs in async jobs whose results are woven into future foreground turns. +- Do not refresh graphify-out unless explicitly requested. + +## Risks +- Sub-300ms end-to-end is likely impossible for full STT -> LLM -> TTS answer when using cloud STT and a cloud LLM; design must measure p50/p95 and separate acknowledgement latency from answer latency. +- Current browser ScriptProcessor uses 4096-frame chunks, which is roughly 85ms at 48kHz before network; AudioWorklet with 20ms frames is needed for tighter latency. +- Current Miso wrapper serializes a full WAV before response; streaming Mimi/audio chunks and cancellation must be added before it can serve live voice well. +- GPT-5.5 API/tool availability, rate limits, and pricing should be verified again at implementation time. +- Voice cloning and prompt-audio continuation need consent, safety, and storage boundaries. + +## Approval Required +- User approval/API keys before live external provider traffic. +- User purchase/selection of a Vast.ai GPU before remote deployment work. +- Approval before broad rewrites of ChatPanel/server voice code or moving voice orchestration into a new FastAPI service. + +## Work Packets +- P1 external research: Thinking Machines, MisoTTS, Deepgram STT, OpenRouter/GPT models, latency. +- P2 current architecture: Graphify-routed inspection of current voice/socket/tool files. +- P3 target architecture: broker protocol, task delegation, caching, latency budget, failure handling. +- P4 implementation readiness: GPU spec, env vars, staged code plan, verification gates. + +## Integration Policy +Accept source-backed facts from official/vendor docs first. Treat marketing latency claims as targets until measured locally. Keep implementation choices compatible with the current LearningAI browser voice UI, local memory/tool ledgers, and beta diagnostics. + +## Verification +- Graphify query performed before source reads. +- Targeted source inspection only. +- Source links recorded in final report. +- Workflow completeness check before handoff. +- No runtime code changes in this research/prep pass. + +## Reusable Artifacts +- final-report.md: architecture decision and implementation-ready plan. +- results/*.md: packet notes for future implementation. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/results/integration.md b/.workflow/sub-300ms-voice-brain-architecture-research/results/integration.md new file mode 100644 index 0000000..af3e3af --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/results/integration.md @@ -0,0 +1,25 @@ +# Integration Result + +## Accepted +- Use a Tutor-owned realtime voice broker with a foreground GPT-4o-mini learner loop and a GPT-5.5 background tool/reasoning loop. +- Preserve current `/api/voice-agent` as fallback while adding the new broker path. +- Treat MisoTTS as TTS-only and implement duplex behavior through always-on STT, cancellation, playback flushing, and background-result insertion. +- Use RTX 4090 24GB or RTX 3090 24GB as the minimal first Vast.ai machine. + +## Rejected +- Do not claim MisoTTS is multi-duplex/full-duplex. +- Do not promise full semantic answer audio under 300ms as a guaranteed end-to-end target with cloud STT plus cloud LLM. +- Do not put GPT-5.5 in the critical foreground response path. + +## Conflicts +- Thinking Machines' native interaction model can handle continuous audio/video internally; our stack is cascaded. Resolution: mimic the interaction shape with a persistent broker and measurement, while being honest about latency. + +## Decisions +- Foreground: browser AudioWorklet -> broker -> Deepgram STT -> OpenRouter GPT-4o-mini -> MisoTTS -> browser. +- Background: broker -> GPT-5.5/tool runner -> result queue -> natural insertion into foreground speech. +- Verification must measure p50/p95/p99 by stage after the GPU is available. + +## Remaining Risks +- Miso streaming/cancellation may require generator-level changes. +- Region/network latency can dominate. +- Live provider API behavior needs key-backed validation. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/results/local-readiness-slice.md b/.workflow/sub-300ms-voice-brain-architecture-research/results/local-readiness-slice.md new file mode 100644 index 0000000..bdce101 --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/results/local-readiness-slice.md @@ -0,0 +1,23 @@ +# Local Readiness Slice + +## Accepted +- Added `/api/voice-broker` beside the existing `/api/voice-agent` path. +- Kept the current Deepgram voice-agent path as fallback. +- Added `VITE_VOICE_BROKER_MODE=custom` as the browser opt-in for the new broker. +- Staged broker metadata for Deepgram STT, OpenRouter GPT-4o-mini foreground teaching, MisoTTS speech, and GPT-5.5 background jobs without requiring provider keys. +- Preserved local brain context handoff through the existing `buildBrainContextPacket` flow. + +## Verification Target +- Local broker accepts `voice_auth`. +- Local broker records active book, previous memory, document context, and proof metadata. +- Local broker handles typed voice injection and stages a GPT-5.5 background job when the user asks for fresh/current/tool work. +- Provider traffic remains deferred until the user supplies keys and the GPU instance. + +## Files Touched +- `server.ts` +- `src/components/ChatPanel.tsx` +- `src/store/index.ts` +- `tests/system-activity.test.mjs` +- `TUTOR_ARCHITECTURE.md` +- `src/lib/userBrainArchitectureBook.ts` +- `src/lib/tutorBook.json` diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/results/vast-gpu-setup.md b/.workflow/sub-300ms-voice-brain-architecture-research/results/vast-gpu-setup.md new file mode 100644 index 0000000..b27ce7f --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/results/vast-gpu-setup.md @@ -0,0 +1,105 @@ +# Vast GPU Setup Result + +## Instance +- SSH: `ssh -p 19691 root@175.155.64.137` +- GPU: NVIDIA GeForce RTX 4090 reported by `nvidia-smi` with 49140 MiB VRAM. +- Driver/CUDA: NVIDIA driver 550.144.03, CUDA runtime 12.4, `nvcc` 12.6. +- Persistent local volume: `/workspace`, 107 GB available. + +## Cache Layout +The container root is only about 32 GB, so model and package caches must stay on +the attached `/workspace` volume: + +```bash +export HF_HOME=/workspace/hf +export HUGGINGFACE_HUB_CACHE=/workspace/hf/hub +export UV_CACHE_DIR=/workspace/uv-cache +export XDG_CACHE_HOME=/workspace/.cache +export MISO_TTS_8B_MODEL=/workspace/hf/hub/models--MisoLabs--MisoTTS/snapshots/ef6b096cc35d3cde6aa0721013648416c14c36b2/model.safetensors +export MISO_TTS_TOKENIZER_MODEL=unsloth/Llama-3.2-1B +``` + +These exports are stored on the instance at: + +```text +/workspace/miso-service/env.sh +``` + +## Miso Install +MisoTTS was cloned to: + +```text +/workspace/MisoTTS +``` + +The repo was synced with: + +```bash +cd /workspace/MisoTTS +source /workspace/miso-service/env.sh +uv sync --python 3.10 +``` + +Tutor's HTTP wrapper was copied to: + +```text +/workspace/miso-service/misotts_api_server.py +``` + +## Service Command +Run from the MisoTTS repo so the wrapper can import Miso's `generator` module: + +```bash +cd /workspace/MisoTTS +source /workspace/miso-service/env.sh +PYTHONPATH=/workspace/MisoTTS uv run python /workspace/miso-service/misotts_api_server.py --host 127.0.0.1 --port 8090 +``` + +Keep the local tunnel open from the Mac: + +```bash +ssh -N -p 19691 root@175.155.64.137 -L 8080:localhost:8090 +``` + +Then configure Tutor with: + +```bash +VITE_VOICE_BROKER_MODE=custom +MISO_TTS_API_URL=http://127.0.0.1:8080 +``` + +## Notes +- First Miso startup downloads the model weights, Mimi codec, SilentCipher, and + tokenizer into `/workspace/hf`. +- Miso's upstream generator defaults to gated `meta-llama/Llama-3.2-1B` for + the tokenizer. Tutor's wrapper supports `MISO_TTS_TOKENIZER_MODEL`; this + instance uses public `unsloth/Llama-3.2-1B`, which reports matching BOS/EOS + ids and vocabulary size. +- The broker is BYOK by default for OpenRouter, Deepgram, and Serper. Shared + server fallback keys require explicit fallback flags. + +## Verification +- Local tunnel health: + +```bash +curl http://127.0.0.1:8080/health +``` + +returned `device: cuda`, `loaded: true`, and tokenizer source +`unsloth/Llama-3.2-1B`. + +- Cold speech smoke after model download: + - status: 200 + - bytes: 7758 + - wall time: 87.77s + - wrapper timing: generator load 80.62s, generation 4.91s + +- Warm speech smoke: + - status: 200 + - bytes: 23118 + - wall time through tunnel: 6.13s + - wrapper timing: generation 3.36s + - output: RIFF WAV, PCM 16-bit mono 24000 Hz + +This proves the GPU service works, but the current wrapper is whole-utterance +TTS. It is not yet a 110ms streaming path. diff --git a/.workflow/sub-300ms-voice-brain-architecture-research/state.json b/.workflow/sub-300ms-voice-brain-architecture-research/state.json new file mode 100644 index 0000000..ffa66d3 --- /dev/null +++ b/.workflow/sub-300ms-voice-brain-architecture-research/state.json @@ -0,0 +1,64 @@ +{ + "title": "Sub 300ms Voice Brain Architecture Research", + "slug": "sub-300ms-voice-brain-architecture-research", + "created_at": "2026-06-07T12:11:03+00:00", + "status": "gpu_setup_verified_latency_caveat", + "approval": { + "required": [ + "external provider keys before live traffic", + "Vast.ai GPU before remote deployment" + ], + "granted": null, + "notes": "Research, architecture context, configured broker adapters, and Vast GPU Miso service setup are verified. Live OpenRouter/Deepgram/Serper traffic still waits on user keys. Miso whole-utterance wrapper works but is not yet a 110ms streaming path." + }, + "packets": [ + { + "id": "P1", + "title": "External research", + "status": "completed" + }, + { + "id": "P2", + "title": "Current architecture", + "status": "completed" + }, + { + "id": "P3", + "title": "Target architecture", + "status": "completed" + }, + { + "id": "P4", + "title": "Implementation readiness", + "status": "completed" + }, + { + "id": "P5", + "title": "Local broker readiness slice", + "status": "verified_local" + }, + { + "id": "P6", + "title": "Vast GPU Miso setup", + "status": "verified_latency_caveat" + } + ], + "verification": { + "status": "passed", + "checks": [ + "graphify query completed before source reads", + "targeted source inspection completed", + "external sources checked", + "local /api/voice-broker staged without provider traffic", + "configured broker adapter test passed with mocked OpenRouter, Serper, and Miso", + "focused node tests passed: tests/system-activity.test.mjs tests/brain-context.test.mjs", + "Vast GPU cache layout written under /workspace", + "MisoTTS health verified through local tunnel with CUDA and loaded model", + "MisoTTS cold speech smoke passed in 87.77s", + "MisoTTS warm speech smoke passed in 6.13s through tunnel", + "npm run lint passed", + "npm run build passed", + "git diff --check passed" + ] + } +} diff --git a/docs/voice-mode-miso-latency-book.md b/docs/voice-mode-miso-latency-book.md new file mode 100644 index 0000000..ca81e8b --- /dev/null +++ b/docs/voice-mode-miso-latency-book.md @@ -0,0 +1,280 @@ +# Voice Mode MisoTTS Latency Book + +Date: 2026-06-07 + +## Chapter 1: What We Tried To Build + +The target architecture was a low-latency voice tutor loop: + +- Deepgram handles speech-to-text from the microphone. +- GPT-4o-mini, through OpenRouter, handles the foreground teaching response. +- GPT-5.5 handles asynchronous background work such as web search, PDFs, code analysis, and deeper tool jobs. +- MisoTTS 8B runs on a rented GPU machine and speaks the assistant response. +- FastAPI wraps the GPU-side TTS service. +- Tutor's local server connects the browser, foreground model, background jobs, STT, and TTS through `/api/voice-broker`. + +After the Miso latency failure, the replacement live target became Deepgram on +both speech edges: Nova streaming STT plus Aura streaming TTS. Browser speech +synthesis is kept only as an explicit fallback, and MisoTTS is kept out of the +live first-audio path. + +The product goal was strict: first usable spoken response under 200-300 ms, with a conversational socket open between the learner and the voice system. + +## Chapter 2: GPU Instance And Remote Setup + +The tested Vast instance was reached through: + +```bash +ssh -p 19691 root@175.155.64.137 -L 8080:localhost:8080 +``` + +The actual working local tunnel mapped local port 8080 to the remote Miso wrapper on port 8090: + +```bash +ssh -fN -o ExitOnForwardFailure=yes -p 19691 root@175.155.64.137 -L 8080:localhost:8090 +``` + +Observed machine class: + +- GPU: 1x RTX 4090, 24 GB advertised class, enough VRAM for the public MisoTTS 8B bfloat16 path. +- OS: Ubuntu 24.x class remote image. +- Remote service: `/workspace/miso-service/misotts_api_server.py` +- Miso repo: `/workspace/MisoTTS` +- Hugging Face cache: `/workspace/hf` +- Miso model snapshot: `/workspace/hf/hub/models--MisoLabs--MisoTTS/snapshots/ef6b096cc35d3cde6aa0721013648416c14c36b2/model.safetensors` + +The Meta tokenizer repo was gated, so the wrapper used: + +```bash +MISO_TTS_TOKENIZER_MODEL=unsloth/Llama-3.2-1B +``` + +The tokenizer was checked for compatibility at the practical level needed here: BOS/EOS and vocabulary behavior matched the expected Llama 3.2 tokenizer family. + +## Chapter 3: What We Implemented Locally + +The local app gained a custom voice broker route: + +- WebSocket path: `/api/voice-broker` +- Feature flag: `VITE_VOICE_BROKER_MODE=custom` +- Current live TTS default: `VOICE_BROKER_TTS_MODEL=aura-2-thalia-en` +- Browser fallback flag: `VITE_VOICE_BROKER_BROWSER_TTS=false` by default, `true` only for explicit fallback tests +- Live TTS deadline: `VOICE_BROKER_TTS_DEADLINE_MS=180` +- Optional Miso endpoint: `MISO_TTS_API_URL=http://127.0.0.1:8080` + +The Miso wrapper gained: + +- `/health` +- `/v1/audio/speech` +- `/v1/audio/prewarm` +- disk WAV cache under `MISO_TTS_CACHE_DIR` +- stage timing logs +- tokenizer source override +- `max_audio_length_ms` minimum lowered to 80 ms for probing + +The browser voice path gained: + +- `ConversationText` playback through `window.speechSynthesis` +- immediate cancellation on barge-in +- immediate cancellation on voice stop +- local broker auth field `browserTts: true` +- no fresh Miso generation in the live first-audio path when browser TTS is active + +## Chapter 4: What The Measurements Showed + +The important finding is that the public MisoTTS 8B wrapper is not a true sub-200 ms live TTS path for arbitrary fresh text on this setup. + +Unit note: + +- `s` means seconds. +- `ms` means milliseconds. +- `us` means microseconds. +- `1 ms = 1,000 us`. +- Values below `1.000 ms` are sub-millisecond local measurements. For example, `0.324 ms` is about `324 us`. + +Observed Miso timings: + +| Case | Result | +| ------------------------------- | ------------------------------------------------- | +| Cold load plus first generation | about 89.59 s total; generator load about 83.70 s | +| Warm fresh generation | about 5-6 s in early simple smoke tests | +| Later fresh generation probes | 16.48 s, 18.63 s, 61.44 s, and 198.59 s | +| Cached WAV on GPU box | effectively instant in wrapper logs | +| Cached WAV through SSH tunnel | often around 1.2-2.6 s from the Mac | + +The remote log showed fresh generation waits behind a single generation lock. That explains the user-visible symptom: "not reading at times." It was not just a browser playback bug; the server could be waiting on whole-utterance generation. + +The issue was therefore two separate latency problems: + +| Issue | What we observed | Why it matters | +| -------------------------------------------- | ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | +| Fresh Miso generation was seconds to minutes | 16.48 s, 18.63 s, 61.44 s, and 198.59 s fresh generations | This cannot satisfy a 200 ms first-audio target. | +| The Miso wrapper returns whole WAV files | Audio is emitted only after token generation, decode, watermark, resample, and WAV encoding | The browser gets no early audio chunk to play. | +| A single generation lock serialized work | Later calls waited behind slow calls | The system can appear to stop reading. | +| SSH tunnel and network path added delay | Cached remote responses were instant on the GPU box but around 1.2-2.6 s from the Mac | Even cached audio was not reliable as the live first-audio path through that tunnel. | +| Cached audio is not the real solution | Cached acknowledgements can be fast, but arbitrary new responses still miss the target | It can hide the problem for "Okay" but not solve tutoring speech. | + +The public Miso source confirmed the architecture: + +1. tokenize text, +2. generate audio tokens frame by frame, +3. decode the full token stack into audio, +4. watermark and resample, +5. encode and return a full WAV. + +There was no exposed streaming API that could emit the first audio chunk before whole-utterance completion. + +## Chapter 5: Thirty Local Latency Trials + +After moving live speech to browser TTS delegation, we ran 30 local websocket trials on the Mac against: + +```text +ws://127.0.0.1:3001/api/voice-broker?language=en +``` + +Each trial opened the broker, sent `voice_auth` with `browserTts: true`, waited for `VoiceBrokerReady`, injected a text turn, and measured: + +- `ConversationText` latency from injection +- `AgentFinishedSpeaking` latency from injection + +No OpenRouter key was used in this latency test, so the foreground reply was the provider-safe local staged response. This measures the local broker loop after removing Miso from the blocking path, not full Deepgram STT plus GPT-4o-mini plus TTS. + +| Trial | ConversationText ms | AgentFinished ms | +| ----- | ------------------: | ---------------: | +| 1 | 0.449 | 0.472 | +| 2 | 0.736 | 0.871 | +| 3 | 0.286 | 0.304 | +| 4 | 0.465 | 0.470 | +| 5 | 0.333 | 0.343 | +| 6 | 0.494 | 0.550 | +| 7 | 0.437 | 0.472 | +| 8 | 0.704 | 0.722 | +| 9 | 0.301 | 0.385 | +| 10 | 0.228 | 0.278 | +| 11 | 0.889 | 0.905 | +| 12 | 0.388 | 0.397 | +| 13 | 0.534 | 0.576 | +| 14 | 0.360 | 0.396 | +| 15 | 0.324 | 0.331 | +| 16 | 0.226 | 0.236 | +| 17 | 0.254 | 0.284 | +| 18 | 0.264 | 0.271 | +| 19 | 0.320 | 0.347 | +| 20 | 0.611 | 0.673 | +| 21 | 0.139 | 0.154 | +| 22 | 0.106 | 0.108 | +| 23 | 0.099 | 0.110 | +| 24 | 0.133 | 0.136 | +| 25 | 0.142 | 0.164 | +| 26 | 0.090 | 0.176 | +| 27 | 0.125 | 0.142 | +| 28 | 0.326 | 0.497 | +| 29 | 0.431 | 0.480 | +| 30 | 0.416 | 0.424 | + +Summary: + +| Metric | ConversationText | AgentFinished | +| ------- | ---------------: | ------------: | +| Count | 30 | 30 | +| Min | 0.090 ms | 0.108 ms | +| p50 | 0.324 ms | 0.347 ms | +| p90 | 0.611 ms | 0.673 ms | +| p95 | 0.704 ms | 0.722 ms | +| Max | 0.889 ms | 0.905 ms | +| Average | 0.354 ms | 0.389 ms | + +The complete 30-trial run took 78.715 ms wall time. + +The local trial values are milliseconds. Because they are below `1.000 ms`, they can also be read as approximate microseconds: + +| Metric | ConversationText | AgentFinished | +| ------ | ---------------------: | ---------------------: | +| p50 | 0.324 ms, about 324 us | 0.347 ms, about 347 us | +| p95 | 0.704 ms, about 704 us | 0.722 ms, about 722 us | +| Max | 0.889 ms, about 889 us | 0.905 ms, about 905 us | + +These numbers are intentionally only the local broker loop after browser TTS delegation. They do not claim the full future stack is sub-millisecond. The full live stack still needs separate measurement for microphone capture, Deepgram STT, GPT-4o-mini response generation, browser speech start, and background job insertion. + +## Chapter 6: Decision + +The decision is to delete the current GPU instance for now and not treat MisoTTS 8B as the live first-audio path. + +MisoTTS 8B can remain useful for: + +- non-realtime read-aloud, +- pre-generated chapter audio, +- cached acknowledgement clips, +- future async high-quality narration, +- experiments if Miso publishes a true streaming or server-optimized path. + +It should not be the live voice loop renderer unless a future model/server proves fresh arbitrary text under the deadline. + +The live voice path should be: + +```text +Browser mic -> Deepgram STT -> Tutor voice broker -> GPT-4o-mini foreground text + -> Deepgram Aura streaming TTS for first speech + -> GPT-5.5 background jobs asynchronously +``` + +Miso can be attached only after it stops blocking the foreground loop. + +The replacement broker defaults are: + +- STT: `nova-3` +- TTS: `aura-2-thalia-en` +- Foreground model: `openai/gpt-4o-mini` through OpenRouter +- Background model: `gpt-5.5` through OpenRouter +- Browser TTS fallback: `VITE_VOICE_BROKER_BROWSER_TTS=true` + +## Chapter 7: Possible Fixes + +The practical fix we applied locally was architectural: + +1. Keep `/api/voice-broker` as the realtime coordinator. +2. Send assistant `ConversationText` immediately. +3. Stream that text through a per-conversation Deepgram Aura TTS websocket. +4. Keep MisoTTS out of the live first-audio path. +5. Keep browser speech as an explicit fallback only. +6. Keep MisoTTS for read-aloud, cached clips, or async narration only. +7. Record the latency and decision in this book so the failed GPU path is not repeated blindly. + +Possible future fixes if we want a better voice than browser TTS: + +| Fix candidate | What it would need to prove | Risk | +| ----------------------------------- | --------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | +| True streaming TTS model | First audio chunk under 200 ms for uncached arbitrary text, with stable p95 | Quality or deployment complexity may be worse. | +| Managed realtime TTS API | Low first-byte audio latency from our region, barge-in/cancel support, predictable cost | Vendor dependency and API cost. | +| Smaller local TTS model | Runs on smaller GPU/CPU and streams chunks immediately | Voice quality may be weaker than Miso. | +| Prewarm/cache common phrases | Fast acknowledgements and common tutor phrases | Does not solve arbitrary fresh responses. | +| Rewrite Miso server for streaming | Decode and send audio chunks before whole utterance completion | Public Miso path did not expose this; may require deep model/vocoder work and still may not hit p95. | +| Move app/server/GPU closer together | Reduce network and SSH tunnel overhead | Does not fix multi-second fresh generation. | + +The main issue to fix is not "make the HTTP route faster." The main issue is that whole-utterance Miso generation is too slow and non-streaming for realtime tutoring. The replacement must stream early audio, cancel cleanly, and keep the foreground websocket free. + +## Chapter 8: What To Look For Next + +The replacement live TTS model or service must prove: + +- first audio under 200 ms for uncached arbitrary text, +- streaming output rather than whole-WAV completion, +- cancellation/barge-in support, +- stable p95 latency, not just one fast demo, +- no single global generation lock that blocks all sessions, +- acceptable voice quality for tutoring, +- predictable deployment on a small GPU or CPU instance. + +Good candidates to test next are true streaming TTS engines or managed realtime speech APIs. A local model is only acceptable if it can stream chunks quickly on the target hardware. + +## Chapter 9: What Remains Local + +The local repo now contains the important reusable work: + +- Tutor voice broker route in `server.ts` +- Deepgram Aura TTS routing plus explicit browser fallback handling in `src/components/ChatPanel.tsx` +- Miso wrapper script in `scripts/misotts_api_server.py` +- environment flags in `.env.example` +- architecture notes in this book + +Deleting the Vast instance loses the remote cache/model files and running service, but not the local integration work. diff --git a/output/pdf/voice-mode-miso-latency-report.pdf b/output/pdf/voice-mode-miso-latency-report.pdf new file mode 100644 index 0000000..2a8cf5d Binary files /dev/null and b/output/pdf/voice-mode-miso-latency-report.pdf differ