Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 23 additions & 3 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -16,11 +16,31 @@ ALLOW_SERVER_DEEPGRAM_FALLBACK=false
SERPER_API_KEY=
ALLOW_SERVER_SERPER_FALLBACK=false

# Voice broker mode:
# - deepgram: Deepgram Voice Agent path.
# - custom: browser -> local broker -> OpenRouter foreground/background models.
# Voice mode is now chosen at runtime in Settings (no rebuild needed):
# - deepgram-duplex (default): the two-model "interaction" mimic. A fast
# foreground model keeps the conversation moving while a background model does
# the heavy lifting. Runs inside this app's own dev/start server — no separate
# host — and needs only a Deepgram key plus an LLM key (OpenAI or OpenRouter;
# the broker auto-detects the provider from the key shape).
# - deepgram-agent: the single Deepgram Voice Agent path.
# - openai-realtime: test/comparison-only speech-to-speech over WebRTC.
# VITE_VOICE_BROKER_MODE below is only the legacy initial default for the
# runtime setting ("deepgram" -> deepgram-agent, anything else -> duplex).
VITE_VOICE_BROKER_MODE=deepgram

# Real OpenAI key. Used by the read-aloud gpt-4o-mini-tts route and, when the
# fallback flag is enabled, to mint OpenAI Realtime client secrets server-side
# for the test voice mode. Realtime is BYOK-first: the browser can send its own
# OpenAI key, so this server key is optional.
OPENAI_API_KEY=
ALLOW_SERVER_OPENAI_FALLBACK=false
# OpenAI Realtime model for the test/comparison voice mode (browser WebRTC).
VITE_OPENAI_REALTIME_MODEL=gpt-realtime

# Set to false on deployed hosts to ignore the client-supplied x-miso-tts-api-url
# header and use only MISO_TTS_API_URL below (blocks localhost port probing).
MISO_TTS_ALLOW_HEADER_URL=true

# Remote voice/signaling server. Leave VITE_VOICE_WS_URL empty to use the same
# host the app is served from. When the app host cannot serve WebSockets
# (e.g. Vercel), point this at the dedicated voice server, e.g.
Expand Down
19 changes: 0 additions & 19 deletions .eslintrc.json

This file was deleted.

139 changes: 84 additions & 55 deletions BUG_REPORT.md
Original file line number Diff line number Diff line change
@@ -1,68 +1,97 @@
# Bug Review Report
# Audit Report

Review of the Tutor-System codebase (server, client memory engines, views,
and the voice/bilateral architecture). Typecheck passes; 277/281 tests pass
(the 3 failures are environmental — `pymupdf` is not installed in the review
container — not code defects).
Full audit of the Tutor-System codebase (server, client memory engines, views,
PDF pipeline, Mermaid rendering, UI, and the voice architecture) plus the
product changes requested alongside it. Verification gate after the work:
`format:check`, `lint` (tsc), `test` (280 node pass / 1 skipped, 591 DOM pass),
and `build` all green, with `pymupdf`/`pymupdf4llm` installed.

## Fixed in this branch
## Security — fixed

### 1. Chat SSE stream was buffered by the compression middleware (high impact)
`server.ts` applies `app.use(compression())` globally, and `/api/chat` streams
Server-Sent Events without flushing. `compression()` gzip-buffers a
`text/event-stream` body and releases it only when the response ends, so the
token-by-token streaming in `ChatPanel` never runs — the user sees a spinner,
then the whole answer at once. Verified in isolation: five events arrived in a
single trailing chunk.
1. **Cross-site WebSocket hijack on `/api/voice-agent` (High).** The Deepgram
Voice Agent upgrade accepted any origin, unlike the custom broker. It now
uses the same `isAuthorizedLocalVoiceBrokerRequest` (same-origin / debug
token) gate, closing the hijack + Deepgram-credit-drain vector.
2. **Mermaid XSS via `securityLevel: "loose"` (High).** Chat diagrams are
influenced by untrusted PDF/OCR/web content and were rendered with HTML
sanitization disabled + `innerHTML`. Switched to `securityLevel: "strict"`
(matching RevisionView), which DOMPurifies the SVG and blocks in-diagram
`click`/`javascript:` directives.
3. **Unbounded per-user SQLite handles (High, DoS + FD/memory leak).** Each
client-supplied `userId` opened a WAL handle that was cached forever.
`LearnerStore` now evicts idle handles with an LRU cap
(`LEARNINGAI_MAX_OPEN_DBS`, default 64).
4. **Document-read identity mismatch (High, best-effort).** `/documents/:id/file`
and `/text` trusted `req.query.userId`. They now reject a request whose
header identity and query identity explicitly disagree, while still allowing
the header-less react-pdf fetch. (True tenant isolation still requires real
auth, which the local-first model defers.)
5. **Oversized-body DoS (Med).** `express.urlencoded` limit lowered from 100mb
to 1mb.
6. **MisoTTS localhost probing (Low).** `MISO_TTS_ALLOW_HEADER_URL=false` lets
deployments ignore the client-supplied Miso URL header.

**Fix:** mark the SSE response `Cache-Control: no-cache, no-transform`
(which `compression` respects and skips), set `X-Accel-Buffering: no`, and call
`res.flush()` after each event write. After the fix, events stream ~200ms apart.
## Memory leaks — fixed

### 2. Full-duplex voice broker was unreachable on a deployed host (high impact)
The `/api/voice-broker` WebSocket upgrade was gated by
`isAuthorizedLocalVoiceBrokerRequest`, which only passed for loopback or a debug
token. On a deployed persistent Node host the broker returned `403`, so the
foreground/background duplex voice design could never run in production.
7. **Hot-mic on unmount.** A live voice session (mic `MediaStream`,
`AudioContext`, `ScriptProcessorNode`, WebSocket) was never released if
`ChatPanel` unmounted mid-session. Added an unmount cleanup.
8. **Web-search cache never evicted.** Expired entries are now dropped and the
map is size-capped.
9. See security #3 (SQLite handle eviction).

**Fix:** also accept the upgrade when the request `Origin` matches the app host
(same-origin, with `x-forwarded-host` support behind a proxy). This blocks
cross-site WebSocket hijacking while letting the deployed app's own page connect.
Keys remain BYOK per connection; server fallback keys stay gated behind the
existing `ALLOW_SERVER_*` env flags. Verified with real handshakes under
`NODE_ENV=production`: same-origin → `101`, cross-site → `403`, no-origin → `403`.
## Correctness — fixed (previously open findings #3–#5)

## Open findings (not yet fixed)
10. **Blank reply when the tool budget is exhausted mid-tool-call (#3).** The
`/api/chat` agent loop could exit with tool results appended but never
synthesized. Added a final no-tools synthesis turn; telemetry now reports an
honest model-turn count (fixes the #5 off-by-one).
11. **TTS deadline too tight (#4).** `VOICE_BROKER_TTS_DEADLINE_MS` default
raised from 180ms to 1500ms.
12. **Misleading TTS model (#5).** `/api/tts` now calls the `gpt-4o-mini-tts`
model it reports instead of silently substituting `tts-1`.
13. **Silent $0 cost (#5).** `openRouterCost` matches pricing across provider
prefixes, so non-`openai/` models no longer read as $0.

### 3. Tool loop can terminate with no synthesized answer (medium)
In the `/api/chat` agent loop (`while (iterations < MAX_ITERATIONS)`), if the
model emits tool calls on the final allowed iteration, the tool results are
appended to the message list but the loop exits before another model turn turns
them into text. `finalContent` for that turn is empty, so the user can get a
blank/truncated reply when the tool-iteration budget is hit. Consider one extra
synthesis turn after the loop when the last turn was a tool call.
## Product changes shipped alongside the audit

### 4. Voice broker TTS deadline is unrealistically tight (medium)
`VOICE_BROKER_TTS_DEADLINE_MS` defaults to `180`ms. MisoTTS/Deepgram audio
rarely returns that fast, so synthesis is almost always aborted before it
arrives. A value around `1500`ms is more realistic.
- **PDF page-aware context.** Extraction switched to page-indexed markdown
(`pymupdf4llm` `page_chunks`); the current reader page's text is now injected
as a labeled `CURRENT PAGE N of M` block (plus adjacent pages) into the tutor
prompt. Previously the model was never told which page the learner was on and
only ever saw the document's opening characters.
- **Mermaid redesign.** Diagrams render only once the source parses cleanly (no
more raw parser errors flashing mid-stream), are static and fully visible
(removed the perpetual dimming + auto-panning "focus tour"), with click-to-
focus zoom, higher contrast, and reduced-motion support.
- **AdminView dark theme.** Converted the light cream/serif diagnostics surface
to the dark Cosmic Obsidian theme for cross-view consistency (RevisionView's
intentional paper look is preserved).
- **Voice: friction-free duplex + Realtime test mode.** Voice mode is a runtime
Setting; the two-model Deepgram duplex is the default and runs in the app's own
server with just a Deepgram key + an OpenAI or OpenRouter key (auto-detected);
a default-off OpenAI Realtime (WebRTC) mode was added for full-duplex
benchmarking with no persistent server.

### 5. Minor
- `iterations + 1` is reported in telemetry after the loop already incremented,
so iteration counts are off by one (`server.ts`).
- `openRouterCost` only falls back on an `openai/` prefix mismatch, so a
non-OpenAI model with a pricing-key mismatch silently reads as `$0`.
## Deferred / not changed (with rationale)

## Deployment note
- **StatusBadge light palette.** Its wrapper is unused in-app (only its
`currentColor` icons are used, which adapt to dark), and its classes are
pinned by tests — re-theming would break tests for no app benefit.
- **`.eslintrc.json` removed.** It referenced parsers/plugins that were never
installed; `npm run lint` runs `tsc`, so the config only implied a lint that
never ran.
- **True multi-tenant auth.** Local profiles remain "not real auth" by design;
document-read hardening is best-effort within that model.
- **`new Function` JS runner / arbitrary model-authored code.** Retained behind
the explicit "Run" gate; the passive Mermaid XSS path is now closed.

Vercel serverless cannot host the persistent WebSocket the voice broker needs
(`server/vercel-handler.ts` never calls `attachWebSockets`). Voice mode requires
a long-running Node host. With finding #2 fixed, setting
`VITE_VOICE_BROKER_MODE=custom` at build time enables the full-duplex broker
same-origin on such a host.
## Deployment notes

## Responsiveness

Measured in a real mobile browser (Chromium, 320–390px) across Study, Analytics,
Revision, and Admin: no horizontal overflow, cards stack, nav collapses to
icons, composer and voice blob fit. The app is already substantially responsive.
- Vercel serverless still cannot host the persistent voice-broker WebSocket, so
`deepgram-duplex` / `deepgram-agent` need a long-running Node host (or the
separate voice server via `VITE_VOICE_WS_URL`). The `openai-realtime` test
mode is the one voice path that works on serverless (browser↔OpenAI WebRTC +
the tiny `/api/realtime/token` endpoint).
- A proxy terminating the voice WebSocket must strip inbound `X-Forwarded-Host`
so the same-origin broker check cannot be spoofed.
31 changes: 26 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -202,16 +202,37 @@ to cloud Postgres/object storage later.

## Voice Modes

- `deepgram` mode uses the Deepgram Voice Agent path.
- `custom` mode opens a browser WebSocket to the local broker. The broker sends
foreground teaching to an OpenRouter-compatible `VOICE_FOREGROUND_MODEL` and
delegates web/code/PDF/tool work to `VOICE_BACKGROUND_MODEL`.
Voice mode is a **runtime setting** in Settings (no rebuild required):

- **`deepgram-duplex` (default)** — the two-model "interaction" mimic of
Thinking Machines' interaction models. A fast foreground model
(`VOICE_FOREGROUND_MODEL`) keeps the live conversation moving while an async
background model (`VOICE_BACKGROUND_MODEL`) does web/code/PDF/tool heavy
lifting, whose result is stitched back in as a spoken aside. It runs inside
the same `npm run dev` / `npm start` Node process that serves the app — **no
separate server to spin up** — and needs only a **Deepgram key** (STT + Aura
TTS) plus an **LLM key**. The LLM key can be OpenAI **or** OpenRouter; the
broker auto-detects the provider from the key shape (`sk-or-…` → OpenRouter,
`sk-…` → OpenAI direct), so "just a Deepgram key + a ChatGPT key" works.
- **`deepgram-agent`** — the single Deepgram Voice Agent path.
- **`openai-realtime` (test / comparison only)** — connects the browser
straight to OpenAI's Realtime API over **WebRTC**, for a true full-duplex,
uninterrupted benchmark against the cheaper mimic. It needs **no persistent
server** (only a tiny `/api/realtime/token` endpoint that also runs on Vercel
serverless) and is BYOK-first (the browser's OpenAI key mints a short-lived
ephemeral secret; the standard key never reaches the WebRTC exchange). This is
intentionally **not** the default and is billed at premium realtime rates.

- Background answers are cleaned before insertion so raw markdown such as
`**Apple**` is not read aloud.
- MisoTTS is optional and experimental. The local broker only accepts loopback
Miso URLs such as `http://127.0.0.1:8080`.
Miso URLs such as `http://127.0.0.1:8080`; set `MISO_TTS_ALLOW_HEADER_URL=false`
on deployed hosts to ignore the client-supplied Miso URL header.
- No route should claim a universal sub-200 ms guarantee. Report latency as
measured p50, p95, failure rate, route, provider, region, and hardware.
- Deployments that terminate the voice WebSocket behind a proxy must strip any
inbound `X-Forwarded-Host` header so the same-origin broker check can't be
spoofed.

## Getting Started

Expand Down
20 changes: 18 additions & 2 deletions scripts/classify_and_extract.py
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,7 @@ def main():
"total_pages": total_pages,
"pages_with_text": pages_with_text,
"content": "",
"pages": [],
"images": [],
"vision_page_limit": MAX_VISION_PAGES,
}
Expand All @@ -77,8 +78,23 @@ def main():
try:
with contextlib.redirect_stdout(io.StringIO()), contextlib.redirect_stderr(io.StringIO()):
import pymupdf4llm
md_text = pymupdf4llm.to_markdown(doc)
result["content"] = md_text
# Single-pass, page-indexed extraction. page_chunks=True returns
# one dict per page so we can hand the model the exact page the
# learner is looking at, and it is no slower than the plain
# whole-document call.
page_chunks = pymupdf4llm.to_markdown(doc, page_chunks=True)
pages = []
content_parts = []
for index, chunk in enumerate(page_chunks):
page_text = (chunk.get("text") if isinstance(chunk, dict) else str(chunk)) or ""
meta = chunk.get("metadata", {}) if isinstance(chunk, dict) else {}
page_num = meta.get("page", index) if isinstance(meta, dict) else index
pages.append({"page_num": int(page_num), "text": page_text})
# Keep an explicit page marker in the joined markdown so the
# whole-document excerpt retains page boundaries.
content_parts.append(f"<!-- page {int(page_num) + 1} -->\n{page_text}")
result["pages"] = pages
result["content"] = "\n\n".join(content_parts)
except Exception as e:
doc.close()
result["error"] = f"Failed to extract markdown: {str(e)}"
Expand Down
Loading