Claude/repo audit voice UI ijn96b - #8
Merged
Merged
Conversation
…context Security: - Gate /api/voice-agent WS upgrade with the same same-origin check as the custom broker (closes cross-site WebSocket hijack / Deepgram credit drain). - Reject document reads whose header identity and ?userId= query disagree. - Lower express.urlencoded limit from 100mb to 1mb (memory-DoS amplifier). - Add MISO_TTS_ALLOW_HEADER_URL switch to ignore client-supplied Miso URL. Memory leaks: - Tear down the voice session (mic/AudioContext/WebSocket) on ChatPanel unmount. - Bound open per-user SQLite handles with an LRU (was unbounded/never closed). - Evict expired/oversized web-search cache entries. Correctness: - Add a final synthesis turn when the tool-iteration budget is exhausted mid-tool-call, so the user no longer gets a blank reply; report honest model-turn counts. - Raise voice broker TTS deadline default 180ms -> 1500ms. - Call the reported gpt-4o-mini-tts model instead of silently substituting tts-1. - Broaden OpenRouter pricing key matching so non-openai models aren't $0. PDF page-aware context: - Extract page-indexed markdown via pymupdf4llm page_chunks (single pass). - Persist a per-page text map and add readDocumentPageText(). - Send the current reader page to /api/chat and inject a labeled "CURRENT PAGE N of M" block plus adjacent pages into the tutor prompt. Remove dead .eslintrc.json (no eslint installed; lint script runs tsc). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
Mermaid (chat flowcharts): - Render only once the diagram source parses cleanly, so half-streamed ```mermaid fences no longer flash raw parser errors; show a quiet skeleton. - Remove the perpetual node dimming + auto-panning "focus tour" that greyed out and constantly moved the diagram; render it static and fully visible. - Make zoom optional (click / Enter to focus a node, "Reset view" to restore), honoring prefers-reduced-motion. - Switch securityLevel from "loose" to "strict" (diagram source is influenced by untrusted PDF/web content) and lift node/edge/text contrast; align font. AdminView: - Convert the light cream/serif diagnostics surface to the dark Cosmic Obsidian theme (surfaces, text, borders) to match Chat/Study/Analytics, while leaving RevisionView's intentional paper look untouched. StatusBadge is intentionally left as-is: its wrapper is unused in-app (only its currentColor icons are used, which adapt to dark) and its palette is pinned by tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
…ealtime test mode Deepgram two-model duplex is now the forefront and runs with no separate server: - Make voice mode a runtime Setting (deepgram-duplex default, deepgram-agent, openai-realtime) instead of the build-time VITE_VOICE_BROKER_MODE, so switching paths needs no rebuild. - Let the duplex broker's foreground/background LLM calls run on a raw OpenAI key or OpenRouter, auto-detected from the key shape (sk-or- vs sk-), so the two-model mimic works with just a Deepgram key + a ChatGPT key. OpenRouter-only hosted web tool is dropped on the OpenAI-direct path. - Duplex TTS already defaults to Deepgram Aura when a Deepgram key is present, so no MisoTTS/extra server is required. OpenAI Realtime (test/comparison only, default off): - New src/lib/realtimeVoice.ts connects the browser straight to OpenAI over WebRTC (full-duplex audio, data-channel tools, transcripts). Needs no persistent server. - New /api/realtime/token endpoint mints a short-lived ephemeral secret (BYOK-first; server key only with ALLOW_SERVER_OPENAI_FALLBACK). Runs on Vercel serverless too; the standard key never reaches the browser. - ChatPanel branches into the realtime path with study-context instructions, a representative tool set (look_at_study_context + web_search), and full teardown on stop/unmount. Docs: README voice section + .env.example updated for the runtime modes and new OPENAI_API_KEY / ALLOW_SERVER_OPENAI_FALLBACK / VITE_OPENAI_REALTIME_MODEL / MISO_TTS_ALLOW_HEADER_URL variables. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
…mode select - BUG_REPORT.md rewritten as the full audit deliverable (security, memory, correctness fixes + product changes + deferred items with rationale). - rendered-settings.test.tsx now selects each combobox by an option it owns, so it is robust to the newly added voice-mode picker rather than pinned to a brittle index. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
Pull request overview
This PR audits and hardens the Tutor system’s voice + chat surfaces, adding a runtime-selectable voice mode (including an OpenAI Realtime WebRTC test path) while improving security, resource management, and prompt/context quality (page-aware PDF context, safer Mermaid rendering).
Changes:
- Introduces runtime
voiceModepersisted in localStorage and exposes it in Settings, with ChatPanel wiring for Deepgram duplex/agent vs OpenAI Realtime (WebRTC). - Adds OpenAI Realtime support (client WebRTC session + server ephemeral token mint endpoint) and improves voice/session cleanup behavior.
- Hardens server reliability/security: Mermaid strict sanitization, bounded caches/DB handles, reduced body-parser limits, identity mismatch rejection for document reads, and more accurate tool-loop telemetry.
Reviewed changes
Copilot reviewed 13 out of 14 changed files in this pull request and generated 5 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/rendered-settings.test.tsx | Makes settings UI test more robust by selecting the intended <select> via unique option values rather than combobox index. |
| src/store/index.ts | Adds voiceMode runtime setting (type + normalization + persistence). |
| src/lib/realtimeVoice.ts | New OpenAI Realtime WebRTC session helper (ephemeral token + RTCPeerConnection + tool-call plumbing). |
| src/components/SettingsModal.tsx | Adds Voice mode picker UI wired to store. |
| src/components/ChatPanel.tsx | Wires runtime voice mode selection, adds OpenAI Realtime start/stop path, improves Mermaid security + rendering behavior, and adds page-aware context fields. |
| server/web-search.ts | Caps and prunes in-memory web-search cache to prevent unbounded growth. |
| server/learner-store.ts | Adds LRU eviction for per-user SQLite handles and persists per-page extracted text; adds page-text read helper. |
| server.ts | Adds OpenAI Realtime token mint endpoint, strengthens document user-id resolution, reduces urlencoded body limit, improves voice broker provider resolution, fixes tool-loop synthesis edge case + telemetry, and tightens WebSocket upgrade authorization. |
| scripts/classify_and_extract.py | Emits page-indexed extraction output (pages) alongside full markdown content. |
| README.md | Updates documentation to reflect runtime voice modes, Realtime test mode, and new deployment/security notes. |
| BUG_REPORT.md | Updates audit report content and summarizes fixed items + shipped product changes. |
| .eslintrc.json | Removes unused ESLint config file. |
| .env.example | Documents runtime voice mode behavior and adds new OpenAI Realtime / Miso header control env vars. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+92
to
+96
| try { | ||
| dc.close(); | ||
| } catch { | ||
| /* ignore */ | ||
| } |
Comment on lines
+125
to
+132
| micStream = await navigator.mediaDevices.getUserMedia({ | ||
| audio: { | ||
| echoCancellation: true, | ||
| noiseSuppression: true, | ||
| autoGainControl: true, | ||
| }, | ||
| }); | ||
| micStream.getTracks().forEach((track) => pc.addTrack(track, micStream!)); |
Comment on lines
+1159
to
+1164
| const allowServerOpenAi = /^(1|true|yes|on)$/i.test( | ||
| String(process.env.ALLOW_SERVER_OPENAI_FALLBACK || "").trim(), | ||
| ); | ||
| const apiKey = | ||
| byokKey || | ||
| (allowServerOpenAi ? sanitizeApiKey(process.env.OPENAI_API_KEY) : ""); |
Comment on lines
+1058
to
+1063
| <label className="text-sm font-medium text-zinc-300 flex items-center gap-2"> | ||
| <Mic size={14} className="text-zinc-400" /> | ||
| Voice mode | ||
| </label> | ||
| <select | ||
| value={voiceMode} |
Comment on lines
+151
to
+154
| private static readonly MAX_OPEN_DATABASES = Math.max( | ||
| 4, | ||
| Number(process.env.LEARNINGAI_MAX_OPEN_DBS || 64), | ||
| ); |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.