Skip to content

Claude/repo audit voice UI ijn96b - #8

Merged
MohamedFuad16 merged 4 commits into
mainfrom
claude/repo-audit-voice-ui-ijn96b
Jul 23, 2026
Merged

MohamedFuad16 merged 4 commits into
mainfrom
claude/repo-audit-voice-ui-ijn96b

Conversation

@MohamedFuad16

Copy link
Copy Markdown
Owner

No description provided.

claude added 4 commits July 23, 2026 11:42
…context

Security:
- Gate /api/voice-agent WS upgrade with the same same-origin check as the
  custom broker (closes cross-site WebSocket hijack / Deepgram credit drain).
- Reject document reads whose header identity and ?userId= query disagree.
- Lower express.urlencoded limit from 100mb to 1mb (memory-DoS amplifier).
- Add MISO_TTS_ALLOW_HEADER_URL switch to ignore client-supplied Miso URL.

Memory leaks:
- Tear down the voice session (mic/AudioContext/WebSocket) on ChatPanel unmount.
- Bound open per-user SQLite handles with an LRU (was unbounded/never closed).
- Evict expired/oversized web-search cache entries.

Correctness:
- Add a final synthesis turn when the tool-iteration budget is exhausted
  mid-tool-call, so the user no longer gets a blank reply; report honest
  model-turn counts.
- Raise voice broker TTS deadline default 180ms -> 1500ms.
- Call the reported gpt-4o-mini-tts model instead of silently substituting tts-1.
- Broaden OpenRouter pricing key matching so non-openai models aren't $0.

PDF page-aware context:
- Extract page-indexed markdown via pymupdf4llm page_chunks (single pass).
- Persist a per-page text map and add readDocumentPageText().
- Send the current reader page to /api/chat and inject a labeled
  "CURRENT PAGE N of M" block plus adjacent pages into the tutor prompt.

Remove dead .eslintrc.json (no eslint installed; lint script runs tsc).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
Mermaid (chat flowcharts):
- Render only once the diagram source parses cleanly, so half-streamed
  ```mermaid fences no longer flash raw parser errors; show a quiet skeleton.
- Remove the perpetual node dimming + auto-panning "focus tour" that greyed
  out and constantly moved the diagram; render it static and fully visible.
- Make zoom optional (click / Enter to focus a node, "Reset view" to restore),
  honoring prefers-reduced-motion.
- Switch securityLevel from "loose" to "strict" (diagram source is influenced
  by untrusted PDF/web content) and lift node/edge/text contrast; align font.

AdminView:
- Convert the light cream/serif diagnostics surface to the dark Cosmic
  Obsidian theme (surfaces, text, borders) to match Chat/Study/Analytics,
  while leaving RevisionView's intentional paper look untouched.

StatusBadge is intentionally left as-is: its wrapper is unused in-app (only
its currentColor icons are used, which adapt to dark) and its palette is
pinned by tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
…ealtime test mode

Deepgram two-model duplex is now the forefront and runs with no separate server:
- Make voice mode a runtime Setting (deepgram-duplex default, deepgram-agent,
  openai-realtime) instead of the build-time VITE_VOICE_BROKER_MODE, so
  switching paths needs no rebuild.
- Let the duplex broker's foreground/background LLM calls run on a raw OpenAI
  key or OpenRouter, auto-detected from the key shape (sk-or- vs sk-), so the
  two-model mimic works with just a Deepgram key + a ChatGPT key. OpenRouter-only
  hosted web tool is dropped on the OpenAI-direct path.
- Duplex TTS already defaults to Deepgram Aura when a Deepgram key is present,
  so no MisoTTS/extra server is required.

OpenAI Realtime (test/comparison only, default off):
- New src/lib/realtimeVoice.ts connects the browser straight to OpenAI over
  WebRTC (full-duplex audio, data-channel tools, transcripts). Needs no
  persistent server.
- New /api/realtime/token endpoint mints a short-lived ephemeral secret
  (BYOK-first; server key only with ALLOW_SERVER_OPENAI_FALLBACK). Runs on
  Vercel serverless too; the standard key never reaches the browser.
- ChatPanel branches into the realtime path with study-context instructions,
  a representative tool set (look_at_study_context + web_search), and full
  teardown on stop/unmount.

Docs: README voice section + .env.example updated for the runtime modes and new
OPENAI_API_KEY / ALLOW_SERVER_OPENAI_FALLBACK / VITE_OPENAI_REALTIME_MODEL /
MISO_TTS_ALLOW_HEADER_URL variables.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
…mode select

- BUG_REPORT.md rewritten as the full audit deliverable (security, memory,
  correctness fixes + product changes + deferred items with rationale).
- rendered-settings.test.tsx now selects each combobox by an option it owns,
  so it is robust to the newly added voice-mode picker rather than pinned to a
  brittle index.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vqigtqdnq6RkLrGp4VNBBd
Copilot AI review requested due to automatic review settings July 23, 2026 12:15
@vercel

vercel Bot commented Jul 23, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
tutor-system-architecture Building Building Preview, Comment Jul 23, 2026 12:15pm

@MohamedFuad16
MohamedFuad16 merged commit 6a3ec76 into main Jul 23, 2026
1 of 2 checks passed
@MohamedFuad16
MohamedFuad16 deleted the claude/repo-audit-voice-ui-ijn96b branch July 23, 2026 12:15

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR audits and hardens the Tutor system’s voice + chat surfaces, adding a runtime-selectable voice mode (including an OpenAI Realtime WebRTC test path) while improving security, resource management, and prompt/context quality (page-aware PDF context, safer Mermaid rendering).

Changes:

  • Introduces runtime voiceMode persisted in localStorage and exposes it in Settings, with ChatPanel wiring for Deepgram duplex/agent vs OpenAI Realtime (WebRTC).
  • Adds OpenAI Realtime support (client WebRTC session + server ephemeral token mint endpoint) and improves voice/session cleanup behavior.
  • Hardens server reliability/security: Mermaid strict sanitization, bounded caches/DB handles, reduced body-parser limits, identity mismatch rejection for document reads, and more accurate tool-loop telemetry.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 5 comments.

Show a summary per file
File Description
tests/rendered-settings.test.tsx Makes settings UI test more robust by selecting the intended <select> via unique option values rather than combobox index.
src/store/index.ts Adds voiceMode runtime setting (type + normalization + persistence).
src/lib/realtimeVoice.ts New OpenAI Realtime WebRTC session helper (ephemeral token + RTCPeerConnection + tool-call plumbing).
src/components/SettingsModal.tsx Adds Voice mode picker UI wired to store.
src/components/ChatPanel.tsx Wires runtime voice mode selection, adds OpenAI Realtime start/stop path, improves Mermaid security + rendering behavior, and adds page-aware context fields.
server/web-search.ts Caps and prunes in-memory web-search cache to prevent unbounded growth.
server/learner-store.ts Adds LRU eviction for per-user SQLite handles and persists per-page extracted text; adds page-text read helper.
server.ts Adds OpenAI Realtime token mint endpoint, strengthens document user-id resolution, reduces urlencoded body limit, improves voice broker provider resolution, fixes tool-loop synthesis edge case + telemetry, and tightens WebSocket upgrade authorization.
scripts/classify_and_extract.py Emits page-indexed extraction output (pages) alongside full markdown content.
README.md Updates documentation to reflect runtime voice modes, Realtime test mode, and new deployment/security notes.
BUG_REPORT.md Updates audit report content and summarizes fixed items + shipped product changes.
.eslintrc.json Removes unused ESLint config file.
.env.example Documents runtime voice mode behavior and adds new OpenAI Realtime / Miso header control env vars.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/lib/realtimeVoice.ts
Comment on lines +92 to +96
try {
dc.close();
} catch {
/* ignore */
}
Comment thread src/lib/realtimeVoice.ts
Comment on lines +125 to +132
micStream = await navigator.mediaDevices.getUserMedia({
audio: {
echoCancellation: true,
noiseSuppression: true,
autoGainControl: true,
},
});
micStream.getTracks().forEach((track) => pc.addTrack(track, micStream!));
Comment thread server.ts
Comment on lines +1159 to +1164
const allowServerOpenAi = /^(1|true|yes|on)$/i.test(
String(process.env.ALLOW_SERVER_OPENAI_FALLBACK || "").trim(),
);
const apiKey =
byokKey ||
(allowServerOpenAi ? sanitizeApiKey(process.env.OPENAI_API_KEY) : "");
Comment on lines +1058 to +1063
<label className="text-sm font-medium text-zinc-300 flex items-center gap-2">
<Mic size={14} className="text-zinc-400" />
Voice mode
</label>
<select
value={voiceMode}
Comment thread server/learner-store.ts
Comment on lines +151 to +154
private static readonly MAX_OPEN_DATABASES = Math.max(
4,
Number(process.env.LEARNINGAI_MAX_OPEN_DBS || 64),
);

This branch was successfully deployed

1 active deployment
Preview — 5419afb4 Deployed Jul 23, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants