CodeTutor AI is a browser-based learning system, not a thin chat wrapper. A learner reads a lesson, edits and runs real code in an isolated workspace, proves the result against server-owned completion rules, and receives context-aware teaching help without exposing answer keys or trusting browser-supplied progress.
This document explains the current production design and the code that implements it. For local setup and validation commands, see Development. For product behavior and screenshots, see the README.
The structure follows industry architecture-description practice without claiming formal certification. ISO/IEC/IEEE 42010:2022 informs the separation of stakeholders, concerns, and views; arc42 informs the selective use of context, runtime, deployment, cross-cutting, decision, quality, and risk sections; and C4 informs diagram scope and zoom level. The result starts with a system-context view, opens the primary application containers, places them into the production deployment, then uses focused dynamic views where sequence and trust matter. Code-level detail stays in the adjacent module map rather than crowding the picture.
| Stakeholder | Architectural concerns |
|---|---|
| Learners | Fast, understandable feedback; durable progress; accessible and private learning; no leaked answer keys |
| Curriculum and product owners | Canonical lessons and completion; consistent tutoring behavior; safe anonymous-to-account conversion |
| Engineers and content authors | Clear ownership boundaries, reproducible development, stable extension seams, actionable validation |
| Operators and SRE | Bounded spend and capacity, observable failure, safe deployment, rollback, recoverable background work |
| Security and privacy reviewers | Untrusted-browser containment, arbitrary-code isolation, least privilege, auditable data and AI controls |
The design optimizes for these quality drivers, in order:
- Learning correctness: progress, mastery evidence, and completion come from canonical server-owned rules.
- Execution safety: arbitrary learner code stays outside the API process and inside a resource-bounded, network-disabled runner.
- Teaching value: tutor responses are contextual, useful, bounded, and checked for answer leakage before usage is finalized.
- State integrity and privacy: identity, ownership, quota, shares, and personal state survive reloads without trusting client assertions.
- Operational control: artifacts, spend, capacity, migrations, failure recovery, and rollback remain observable and bounded.
- Product experience: the workspace remains responsive, accessible, and coherent across anonymous, authenticated, responsive, and theme states.
The diagrams deliberately zoom rather than repeat one overloaded picture. The first view treats CodeTutor AI as a single system in its environment; the container view opens that boundary into the primary applications and data dependencies; the deployment view then places those responsibilities onto the production platform. Focused runtime views explain the interactions where ordering, authority, or failure behavior matter. Simple modules remain in tables and prose instead of receiving decorative component diagrams.
| View | Scope and intended reader | Question it answers |
|---|---|---|
| System context | One software system; any product or technical reader | Who uses CodeTutor AI, and which external systems does it depend on? |
| Container architecture | Primary interactive applications, data stores, and providers; product and technical readers | Which runnable parts make up the learning path, and what does each one own? |
| Deployment topology | Production hosting and compute boundaries; engineers and operators | Where do the web, API, and execution responsibilities run? |
| Authentication runtime | One protected-request journey; application and security engineers | Where is identity established and verified? |
| Execution runtime | One code-run journey; backend and security engineers | How does untrusted code reach an isolated runner and return safely? |
| Tutor pipeline | One AI-turn journey; product, AI, backend, and finance owners | Where are context, admission, policy, and spend controlled? |
| Release and rollback | One production change; maintainers and operators | How is an exact candidate validated, promoted, verified, and recovered? |
This is the deliberately small entry diagram. CodeTutor AI is one system here; the container architecture below opens that box.
%%{init: {"flowchart":{"curve":"basis","htmlLabels":true}}}%%
flowchart LR
accTitle: CodeTutor AI system context
accDescr: Shows learners and maintainers using CodeTutor AI and its dependencies on Supabase, OpenAI, and Azure platform services.
learner["Learner<br/>reads, writes, runs, and asks"]
operator["Maintainer<br/>authors, releases, and operates"]
system["CodeTutor AI<br/>browser learning, execution, and tutoring"]
supabase["Supabase<br/>identity, Postgres, and object storage"]
openai["OpenAI<br/>tutor generation"]
azure["Azure platform services<br/>hosting, overflow compute, email, and telemetry"]
learner -->|"learns through"| system
operator -->|"develops and operates"| system
system -->|"stores identity and product state"| supabase
system -->|"requests bounded tutor turns"| openai
system -->|"runs and observes production"| azure
classDef person fill:#eef2ff,stroke:#4f46e5,color:#0f172a,stroke-width:2px;
classDef product fill:#f5f3ff,stroke:#7c3aed,color:#0f172a,stroke-width:3px;
classDef external fill:#f8fafc,stroke:#64748b,color:#0f172a,stroke-width:2px;
class learner,operator person;
class system product;
class supabase,openai,azure external;
Legend: blue boxes are people, purple is the CodeTutor AI system boundary, and gray boxes are external software systems or platforms. Every arrow is a labeled dependency directed from its user to its provider.
This C4-style container view is the primary technical overview of the interactive learning path. A container in this view is a runnable application or data boundary, not necessarily a Docker container. The crawler-facing share adapter and hosting details stay in the deployment view below; secondary operational services stay in the observability section.
%%{init: {"flowchart":{"curve":"basis","htmlLabels":true}}}%%
flowchart LR
accTitle: CodeTutor AI container architecture
accDescr: Shows the primary interactive applications, data boundaries, external providers, and their responsibilities.
learner["Learner"]
operator["Maintainer"]
subgraph system["CodeTutor AI"]
direction LR
web["Web application<br/><b>React, TypeScript, Vite, Monaco</b><br/>lessons, editor, tutor, and progress"]
api["Application API<br/><b>Express and TypeScript</b><br/>authorization, canonical learning state,<br/>execution and tutor orchestration"]
execution["Execution plane<br/><b>ExecutionBackend and isolated runners</b><br/>safe code runs and protected grading"]
end
auth["Supabase Auth<br/>session identity"]
data[("Supabase data plane<br/>Postgres, RLS, and object storage")]
openai["OpenAI Responses API<br/>structured tutor generation"]
learner -->|"uses"| web
operator -->|"administers through"| web
web <-->|"authenticates and refreshes"| auth
web -->|"HTTPS product requests"| api
api -->|"reads and writes canonical state"| data
api -->|"dispatches code and tests;<br/>receives normalized results"| execution
api -->|"requests bounded tutor turns"| openai
classDef person fill:#eef2ff,stroke:#4f46e5,color:#0f172a,stroke-width:2px;
classDef client fill:#e0f2fe,stroke:#0284c7,color:#0f172a,stroke-width:2px;
classDef authority fill:#f5f3ff,stroke:#7c3aed,color:#0f172a,stroke-width:3px;
classDef compute fill:#fff7ed,stroke:#ea580c,color:#0f172a,stroke-width:2px;
classDef dataNode fill:#ecfdf5,stroke:#059669,color:#0f172a,stroke-width:2px;
classDef external fill:#f8fafc,stroke:#64748b,color:#0f172a,stroke-width:2px;
class learner,operator person;
class web client;
class api authority;
class execution compute;
class data dataNode;
class auth,openai external;
Legend: blue boxes are people and browser presentation, purple is the application authority, orange is isolated execution, green is durable product data, and gray boxes are external managed systems. Arrows are labeled runtime request or data paths.
Three boundaries shape the system:
- The browser is an untrusted learning client. It owns interaction state and renders the workspace, but it does not authorize access, award completion, set quota, or define canonical lesson context.
- The API is the application authority. It authenticates requests, resolves curriculum and learner state, meters AI usage, manages execution sessions, and persists product data.
- Learner code runs outside the API process. Each active workspace receives an isolated runner. A socket proxy exposes only the Docker operations the session manager needs; optional Azure Container Instances provide bounded overflow capacity.
This view places the web, API, and execution responsibilities onto the actual production applications and compute planes. Managed-service dependencies stay in the container view so this diagram can focus on hosting and traffic. Arrows describe a request or data relationship; they do not imply shared trust.
%%{init: {"flowchart":{"curve":"basis","htmlLabels":true}}}%%
flowchart TB
accTitle: CodeTutor AI deployment topology
accDescr: Shows the browser, static web edge, authentication, API ingress, application authority, and local or overflow execution planes.
learner["Learner<br/>uses a web browser"]
crawler["Search or social crawler<br/>requests public share metadata"]
swa["Azure Static Web Apps<br/>application and prerendered catalog"]
sharefn["SWA managed Function<br/>crawler share metadata"]
auth["Supabase Auth<br/>browser session identity"]
caddy["Azure VM · Caddy<br/>TLS and API routing"]
api["Azure VM · Express API<br/>application authority and orchestration"]
local["VM execution plane<br/>socket proxy and local runners"]
aci["ACI overflow subnet<br/>runner group and sidecar"]
learner -->|"loads HTML, JS, CSS"| swa
learner <-->|"authenticates and refreshes"| auth
learner -->|"calls HTTPS /api"| caddy
crawler -->|"loads a public share URL"| sharefn
caddy -->|"forwards API traffic"| api
sharefn -->|"HMAC-signed preview request"| caddy
api -->|"runs learner code by default"| local
api -->|"bursts when enabled and at capacity"| aci
classDef client fill:#eef2ff,stroke:#4f46e5,color:#0f172a,stroke-width:2px;
classDef edgeNode fill:#ecfeff,stroke:#0891b2,color:#0f172a,stroke-width:2px;
classDef app fill:#f5f3ff,stroke:#7c3aed,color:#0f172a,stroke-width:2px;
classDef compute fill:#fff7ed,stroke:#ea580c,color:#0f172a,stroke-width:2px;
classDef managed fill:#f8fafc,stroke:#64748b,color:#0f172a,stroke-width:2px;
class learner,crawler client;
class swa,sharefn,caddy edgeNode;
class api app;
class auth managed;
class local,aci compute;
Legend: blue boxes are external requesters; cyan is the public edge/ingress; purple is CodeTutor application authority; orange is isolated execution; and gray is an external identity system. Solid arrows are runtime request/data paths.
| Layer | Owns | Primary implementation |
|---|---|---|
| Web client | Routes, lesson/workspace presentation, Monaco, tutor rendering, optimistic interaction state | frontend/src/App.tsx, frontend/src/api/client.ts, frontend/src/state |
| Static web edge | SPA/static assets, prerendered catalog and lesson pages, crawler share adapter | frontend/vite.config.ts, swa-api/src/sharePage.js |
| Application API | Authentication boundary, canonical state, execution orchestration, AI admission and policy | backend/src/index.ts, backend/src/routes |
| Execution plane | Session ownership, local runner lifecycle, optional ACI overflow, function-test harness | backend/src/services/session, backend/src/services/execution |
| Data plane | Durable learner/product state, RLS, share assets | backend/src/db, supabase/migrations |
| Delivery and operations | Immutable images, candidate validation, promotion, rollback, telemetry | .github/workflows/release.yml, infra/azure, docker-compose.prod.yml |
The production shape is constrained intentionally:
- the learning workspace must run in a standard browser with no local toolchain;
- the public application is static-edge hosted while stateful authority remains in the API;
- Supabase provides identity, Postgres, and object storage, but product authorization still belongs to the API;
- arbitrary code needs Docker-compatible isolation locally and on the VM, with Azure Container Instances as optional overflow;
- tutoring depends on the OpenAI Responses API but remains wrapped in provider-independent admission, context, output-policy, and settlement boundaries; and
- database and release changes must be forward, reviewable, and recoverable.
The strategy follows from those constraints: a React static client, an Express application authority, server-owned learning state, ephemeral per-session runners behind an execution abstraction, structured AI output with atomic reservations, forward Postgres migrations with RLS, and immutable candidate artifacts promoted through reusable gates.
Supabase Auth is the identity provider. The browser uses the public Supabase client only for authentication and session refresh. Product reads and writes go through the API, which validates the bearer token and applies its own authorization and ownership rules.
%%{init: {"sequence":{"mirrorActors":false,"useMaxWidth":true,"wrap":true}}}%%
sequenceDiagram
accTitle: Authentication and application hydration
accDescr: Shows browser authentication, local JWT verification with cached JWKS, canonical state loading, and protected rendering.
autonumber
actor L as Learner
participant F as React frontend
participant A as Supabase Auth
participant B as Express API
participant D as Postgres
L->>F: Sign in or return with a session
F->>A: Authenticate or refresh
A-->>F: Access token
F->>B: Product request with Bearer token
B->>B: Verify JWT with cached JWKS
B->>D: Read canonical user-owned state
D-->>B: Preferences, progress, project, usage
B-->>F: Hydration response
F-->>L: Render protected experience
The route and hydration boundary begins in frontend/src/App.tsx. API calls pass through frontend/src/api/client.ts, which attaches the active bearer token, identifies browser mutations, cancels stale requests, and performs one bounded refresh retry. Server authentication and route composition begin in backend/src/index.ts.
The browser never talks to Docker or ACI directly. It submits a project snapshot or execution request to the API. The server checks session ownership, resolves the assigned backend handle, validates paths and limits, then executes inside the learner's runner.
%%{init: {"sequence":{"mirrorActors":false,"useMaxWidth":true,"wrap":true}}}%%
sequenceDiagram
accTitle: Isolated code execution request
accDescr: Shows a workspace request crossing the API and session manager into an isolated runner before normalized results return.
autonumber
participant W as Workspace UI
participant B as Express API
participant S as Session manager
participant E as Execution backend
participant R as Isolated runner
W->>B: Snapshot and execute
B->>B: Validate request
B->>S: Require owned active session
S->>E: Dispatch by session handle
alt Local capacity available
E->>R: Execute through allowlisted socket proxy
else Bounded overflow admitted
E->>R: Execute through authenticated ACI sidecar
end
Note over R: Non-root, no network, read-only rootfs,<br/>writable session workspace, CPU, memory, PID, and time limits
R->>R: Write files and launch learner process
opt Protected function tests
R->>R: Remove hidden expectations before learner code
R->>R: Trusted parent harness consumes one-time nonce,<br/>then signs the post-exit result
end
R-->>E: stdout, stderr, status, optional test envelope
E-->>B: Normalized execution result
opt Guided completion
B->>B: Verify proof and apply canonical completion rules
end
B-->>W: Output or canonical result
The execution abstraction is defined in backend/src/services/execution/backends/types.ts:
localDocker.tscreates one local runner per active session through the allowlisted socket proxy.aci.tscontrols Azure Container Instances through managed Azure APIs and the authenticated sidecar.hybrid.tskeeps work local until configured capacity or operating rules direct overflow to ACI, then dispatches later calls by the stored handle kind.index.tsalways constructs the local backend first. It returns that backend directly when overflow is disabled or incomplete, and otherwise wraps local plus ACI in the hybrid router.
There are therefore two effective runtime shapes: local-only, and
hybrid local plus ACI overflow. ACI is an implementation inside the hybrid
shape, not a separately selected operating mode. The current factory does not
read config.executionBackend; the legacy EXECUTION_BACKEND setting remains
in configuration but is not an active routing selector. Treating that setting
as operational control would be incorrect until the code either removes it or
wires and validates it explicitly.
Completion rules are curriculum authority, so the browser does not receive private expected values. Public lesson content supplies presentation and safe rule metadata; protected function-test data lives under content/memory-warmups and backend-owned lesson sources.
For function_tests, the runner harness:
- loads expectations into memory and removes the tests file before learner code starts;
- launches learner code in a child process inside the already isolated runner;
- emits a sentinel-wrapped result envelope signed with a per-run HMAC nonce; and
- lets
runHarness.tsverify the signature with a timing-safe comparison before the API trusts the result.
The JavaScript vm context narrows the driver environment but is not treated as the security boundary. The container boundary, process isolation, resource controls, and signed envelope form the trust model.
Tutor turns combine deterministic admission and output policy with model-generated teaching. The browser may submit code, selection, run output, history, and lesson identifiers, but all of those are untrusted evidence. For guided lessons, the backend resolves canonical course content, learner progress, and teaching stage before building the prompt.
%%{init: {"flowchart":{"curve":"basis","htmlLabels":true}}}%%
flowchart TB
accTitle: Contextual tutor request pipeline
accDescr: Shows authority and admission, bounded model generation, output policy, usage settlement, and shared response rendering.
subgraph admission["1 · Establish authority and admit the turn"]
direction LR
request["Tutor request<br/>question plus untrusted workspace evidence"]
authz["Credential and access<br/>platform allowance or learner BYOK"]
context["Canonical context<br/>lesson, progress, evidence, stage"]
evidence["Contextual evidence verification<br/>API-minted run-receipt chain<br/>server-owned episode identity"]
modelPolicy["Server model policy<br/>operator override or compiled fallback"]
route["Model resolution<br/>platform authority or learner BYOK choice"]
request --> authz --> context --> route
request -. "submits prior signed run receipts" .-> evidence
context -. "binds receipts to canonical state" .-> evidence
modelPolicy --> route
end
subgraph generation["2 · Generate bounded teaching output"]
direction LR
reserve["Atomic database admission<br/>quota, spend, and replay claims"]
prompt["Prompt builder<br/>intent, pedagogy, bounded context"]
model["OpenAI Responses API<br/>strict structured output"]
reserve --> prompt --> model
end
subgraph delivery["3 · Validate, settle, and render"]
direction LR
policy["Output policy<br/>value, safety, answer leakage"]
settle["Usage settlement<br/>ledger and cancellation recovery"]
render["Shared tutor renderer<br/>streamed or JSON response"]
policy --> settle --> render
end
route --> reserve
evidence --> reserve
model --> policy
classDef input fill:#eef2ff,stroke:#4f46e5,color:#0f172a,stroke-width:2px;
classDef authority fill:#f5f3ff,stroke:#7c3aed,color:#0f172a,stroke-width:2px;
classDef provider fill:#fff7ed,stroke:#ea580c,color:#0f172a,stroke-width:2px;
classDef result fill:#ecfdf5,stroke:#059669,color:#0f172a,stroke-width:2px;
class request input;
class authz,context,evidence,modelPolicy,route,reserve,prompt,policy,settle authority;
class model provider;
class render result;
Legend: blue is untrusted request evidence, purple is CodeTutor-owned authority or policy, orange is the external model provider, and green is the validated response boundary delivered to the learner.
Key modules:
backend/src/routes/ai.tsis the request and streaming composition boundary.canonicalTutorContext.tsresolves server-authoritative lesson and learner context.contextualTutor.tsresolves a contextual offer from canonical lesson state, current files, and the latest server result.contextualEvidence.tsverifies the complete signed run-evidence chain and derives the server-owned error-episode identity.tutorIntent.ts,tutorProgress.ts, and the prompt builders turn that context into teaching instructions.modelRouting.tsowns the single compiled platform fallback.platformTutorModel.tsresolves an optional auditedsystem_configoverride and fails safely to that fallback when configuration is absent, invalid, unpriced, or temporarily unreadable. Platform-funded browser requests do not carry a model choice.modelRegistry.tsdefines Tutor compatibility and records which models completed the teaching-quality gate. The admin and BYOK pickers discover compatible GPT-5+ text models from the relevant OpenAI key; Luna is ranked first as the recommended evaluated choice. Specialized audio, realtime, image, Codex, search, and long-running Pro variants are excluded from the current bounded Tutor request contract.- Platform candidates must also have a backend-owned price entry before activation. A more expensive override requires an explicit cost acknowledgement, a reason, a final confirmation, optimistic concurrency against
set_at, and an append-only admin audit event. Clearing the override is a separately confirmed, audited rollback to the compiled fallback. - BYOK remains a learner-owned credential and model-choice path. Existing retired selections are rehydrated from the key's current compatible model list before the Tutor composer becomes interactive; the server still rejects a BYOK request that omits the learner-selected model.
aiReservations.tsmakes quota and spend admission atomic, claims both the qualifying evidence chain and its error episode against replay, and reconciles abandoned reservations.openaiProvider.tscalls the Responses API with bounded output and structured schemas.tutorOutput.tsandtutorPolicy.tsvalidate usefulness and safety before visible usage is finalized.frontend/src/util/useTutorAsk.tsis the shared client request path;TutorResponseViews.tsxis the shared visual renderer.
Only a turn that passes the product's teaching-value contract counts as a visible question. The execution endpoint mints each signed run receipt from the workspace snapshot and server result captured together under the session workspace lock. A contextual offer cannot then be replayed by changing client-owned epoch or revision fields: the database claims the server-owned error episode and its complete receipt chain inside the bounded evidence window. Reservations fail closed under uncertainty, while cancellation and crash reconciliation prevent abandoned work from silently consuming capacity forever.
Course and lesson progress, editor snapshots, preferences, saved tutor messages, streak history, concept evidence, usage, and public-share ownership are durable server-side state. Frontend stores provide responsive local interaction but reconcile against the API.
Public shares have two read paths:
- A human opens
/s/:tokenin the React application and reads through the bounded public API. - A crawler reaches the managed SWA function, which signs a purpose-bound request to the backend's internal preview endpoint. That path returns metadata without counting a human view and fails closed if the dedicated HMAC configuration is absent.
Share story and Open Graph images live in Supabase Storage. Durable schema and policy truth lives exclusively in ordered forward migrations under supabase/migrations.
The frontend is React, TypeScript, Vite, React Router, Zustand, Monaco, and Tailwind with semantic design tokens.
- Public product: landing, comparison, privacy, terms, support, login/signup/reset, auth callback, and public share.
- Anonymous trial:
/try/lesson/:courseId/:lessonId; the product boundary limits anonymous learning to the supported first-lesson experience even if a different URL is supplied. - Authenticated learning: start/welcome, free editor, course library, saved tutor messages, course and lesson workspaces.
- Internal: development content health and authenticated admin routes, each guarded at the route and server boundary.
| State | Owner |
|---|---|
| Authentication session | Supabase client plus the application auth boundary |
| Durable preferences, progress, project, shares, saved messages | API/Postgres; frontend stores cache and reconcile |
| Editor buffers and layout interaction | Project and preference stores, synchronized at explicit persistence boundaries |
| Run lifecycle and output | Run/session stores, backed by server execution sessions |
| Tutor availability and quota presentation | useAIStatus cache; backend ledger and reservations remain authoritative |
Course JSON under frontend/public/courses is browser-safe curriculum content. Build-time scripts generate the course registry, prerendered public pages, sitemap, and Open Graph assets. Generated registry caches are not authored sources and should not be committed.
backend/src/index.ts is the Express composition root. Its order is intentional:
- CORS, Helmet, JSON parsing, request identifiers, structured logging, and metrics;
- narrowly scoped public endpoints such as health and email unsubscribe;
- authenticated session, project, execution, AI, user-data, feedback, and share endpoints;
- anonymous trial and anonymous-to-account handoff boundaries;
- admin-only and internal HMAC-protected routes; and
- not-found and error handling.
Route-specific middleware owns body-size limits, CSRF/mutation checks, rate limits, token verification, admin authorization, and response redaction. The route tree in the composition root is the authoritative API inventory; individual route modules define their schemas and ownership rules.
The server begins listening before asynchronous readiness completes so health reporting can distinguish liveness from readiness. Background services start only after the required dependencies are ready. They include session cleanup, cost and capacity sampling, ACI health and warm-pool control, platform-budget monitoring, email digest processing, orphan-share cleanup, invariant validation, AI-reservation reconciliation, and abandoned-progress repair.
The backend uses a privileged server connection where an operation genuinely requires it, including controlled Supabase Auth administration. For user-scoped reads and writes, explicit user_id predicates and RLS-context helpers provide defense in depth. The browser receives only the public Supabase URL and publishable key; the service-role key and database credentials are server-only secrets.
Representative durable domains include:
- user preferences, editor projects, course and lesson progress;
- usage ledgers, atomic AI reservations, contextual evidence and episode claims, overrides, deny lists, and audit events;
- saved tutor messages, streaks, streak days, shares, and view telemetry;
- concept ledger, evidence, retrieval, and memory-warmup state;
- system configuration, release/operations state, and evaluation governance.
Server boundaries own authorization, quota, protected curriculum data, mastery evidence, and completion. A client-provided identifier narrows a request; it never proves ownership or truth.
| Boundary | Invariant |
|---|---|
| Browser to API | Bearer identity is verified server-side; mutations use origin/CSRF defenses, bounded bodies, and route-specific rate limits. |
| Learner project paths | Paths are normalized and allowlisted; traversal, protected-file collision, and symlink escape fail closed. |
| Runner isolation | Non-root process, read-only root filesystem, dropped capabilities, no-new-privileges, disabled network, PID/CPU/memory/time limits. |
| Docker control | API reaches an allowlisted socket proxy, not a broadly exposed Docker socket. |
| Function tests | Expected values are removed before learner code; results require a per-run HMAC envelope. |
| Tutor context | Browser content is untrusted evidence; canonical lesson/progress context is resolved by the server, and contextual offers require a signed run-evidence chain plus a database-owned episode claim. |
| Platform AI | Server-controlled model allowlist, atomic admission, per-user/global caps, deny list, and kill switch. |
| BYOK | Keys are encrypted with AES-256-GCM using a server-held master key and are never returned to the browser. |
| Logging | Project, execution, and AI payloads redact to bounded metadata unless an explicit local-only debugging switch is enabled. |
| Secrets | Production secrets are sourced through Key Vault/managed identity and removed from process.env after validated configuration is built. |
| Account deletion | Live sessions and user-owned product data are removed before the controlled Supabase Auth admin deletion completes. |
The API exposes /api/health for readiness and /api/health/deep for dependency-aware production probes. /api/metrics is loopback-only unless a bearer token is explicitly configured.
Structured logs and metrics feed Azure Monitor, Application Insights, and Log Analytics. The infrastructure defines resource health, CPU, memory, disk, OOM, email-delivery, key-decryption, unhandled-rejection, spend, and availability alerts. Cost controls include bounded log ingestion, resource budgets, platform AI circuit breakers, and ACI daily limits.
These decisions explain the current shape. A decision that needs a longer
history or migration plan belongs in a dedicated ADR; the public-share preview
boundary is documented in
ADR_0A_SHARE_PREVIEW_AUTH.md.
| Decision | Why | Consequence and trade-off |
|---|---|---|
| Treat the browser as untrusted and the API as product authority | Learner state, answer keys, quota, and ownership cannot safely depend on mutable client data | More server round trips and canonical resolvers, in exchange for consistent authorization and progress |
| Give each active coding session an ephemeral isolated runner | Arbitrary learner code must not share the API process or another learner's workspace | Strong containment and cleanup boundaries, with startup and capacity overhead |
| Use local execution first and ACI only as optional hybrid overflow | Local runners are predictable and economical; burst capacity should remain bounded | The VM is the normal capacity ceiling unless both operational and cost gates admit overflow |
| Require structured tutor output plus deterministic policy | Model output must render consistently, teach rather than leak answers, and produce value per charged turn | Additional schemas, policy checks, and eval gates around a nondeterministic provider |
| Reserve platform AI usage atomically before generation | Concurrent requests must not overspend per-user or global allowance | Reservation reconciliation is required after cancellation, timeout, or crash |
| Use ordered forward migrations and RLS defense in depth | Applied database history must remain auditable and user-owned rows need an independent database boundary | App rollbacks cannot reverse schema history; migrations must remain backward compatible or be followed by compensation |
| Promote immutable release candidates by digest and manifest | Validation should cover the exact backend, runner, and web artifacts that reach production | Release metadata and artifact retention become operational dependencies, but rollback is reproducible |
| Risk or debt | Current control | Remaining concern |
|---|---|---|
| Stateful API and local runner capacity concentrate on one production VM | Health probes, resource alerts, session caps, restart-safe cleanup, and optional ACI overflow | A VM outage or saturation still has a wider blast radius than a horizontally replicated authority plane |
| Supabase, OpenAI, Azure control planes, and ACS are managed dependencies | Timeouts, bounded retries, fail-closed admission, readiness signals, and learner-facing recovery | Regional or provider outages can still degrade auth, tutoring, execution overflow, email, or persistence |
| Tutor behavior is nondeterministic | Structured output, deterministic policy, compatibility and pricing gates, a complete quality gate for the recommended model, explicit warnings for unevaluated operator choices, and usage settlement | An operator-selected or provider-updated model can change teaching tone or quality without a code-shape change; quality monitoring remains necessary after override |
| Shared development data is a finite integration resource | Isolated E2E identities, cleanup, sharding discipline, and real-database ownership tests | Parallel local and CI activity can still create contention if fixtures bypass isolation rules |
| ACI overflow adds cold-start, networking, and spend variability | Feature flag, runtime switch, capacity cap, daily cost reservation, warm-pool and health controls | It increases operational complexity and is not a substitute for tested local capacity planning |
EXECUTION_BACKEND is configured but not consumed by the factory |
Documentation names the two effective shapes and tests exercise the factory | The unused setting can mislead operators until removed or intentionally wired |
Production promotion validates the exact artifacts that are deployed.
%%{init: {"flowchart":{"curve":"basis","htmlLabels":true}}}%%
flowchart TB
accTitle: Production release and rollback
accDescr: Shows immutable candidate creation, validation, promotion, verification, retention, and explicit rollback of a prior successful candidate.
change["Main branch change"]
scope["Resolve affected release surfaces"]
build["Build backend and runner images"]
ghcr["GHCR immutable digests"]
web["Build frontend and SWA function bundle"]
manifest["Candidate manifest<br/>artifact digests and release metadata"]
gates["Reusable CI, E2E, security,<br/>and contract gates"]
migrations["Conditional migration-state check<br/>backend changes only"]
selected["Selected verified candidate<br/>manifest and immutable artifacts"]
promotevm["VM promotion<br/>Caddy, API, runner digest"]
promoteswa["SWA promotion<br/>static client and share function"]
probes["Production synthetic and health verification"]
retained["Prior successful candidate<br/>manifest, digests, and SWA bundle"]
rollback["Explicit rollback workflow<br/>verify run, SHA, and manifest"]
change --> scope
scope --> build --> ghcr --> manifest
scope --> web --> manifest
manifest --> gates
gates --> migrations --> selected
promotevm --> probes
promoteswa --> probes
probes -. "retain successful candidate" .-> retained
retained --> rollback
rollback --> selected
selected --> promotevm
selected --> promoteswa
classDef source fill:#eef2ff,stroke:#4f46e5,color:#0f172a,stroke-width:2px;
classDef artifact fill:#f5f3ff,stroke:#7c3aed,color:#0f172a,stroke-width:2px;
classDef gate fill:#fff7ed,stroke:#ea580c,color:#0f172a,stroke-width:2px;
classDef deploy fill:#ecfdf5,stroke:#059669,color:#0f172a,stroke-width:2px;
classDef safety fill:#f8fafc,stroke:#64748b,color:#0f172a,stroke-width:2px;
class change,scope source;
class build,ghcr,web,manifest,selected artifact;
class gates,migrations gate;
class promotevm,promoteswa,probes deploy;
class retained,rollback safety;
Legend: blue is source/scope, purple is an immutable artifact, orange is a release gate, green is a production action or verification, and gray is retained rollback evidence or explicit recovery control. Dotted flow is retention; solid flow selects and promotes a verified candidate.
release.yml uses GitHub OIDC for Azure access, builds immutable backend and runner images, creates the SWA candidate, invokes reusable validation workflows against the candidate, verifies database migration state, and promotes only after gates succeed. vm-promote-candidate.sh and the release manifest contracts preserve rollback information rather than relying on mutable tags.
rollback-release.yml accepts only a
successful prior production release run, its full recorded Git SHA, and an
explicit acknowledgement. It verifies that immutable manifest, promotes its VM
image digests and SWA bundle, then probes deployed identity and readiness.
VM promotion can restore its prior candidate when that step fails; promotion
across the VM and SWA is not one distributed transaction. A later surface or
probe failure leaves the release failed for operator assessment and, when
needed, the explicit rollback workflow.
Rollback does not reverse database migrations; forward migrations must stay
compatible with the previous application or be repaired by a new compensating
migration.
Quality is layered deliberately:
- Unit, type, build, content, solution, policy, and contract tests run in CI across the supported host matrix.
- Chromium E2E runs exhaustively in ten empirically selected shards with two workers each; a tracked capacity contract forces remeasurement when suite size leaves the measured band, and metadata-owned critical coverage continues as advisory shadow evidence.
- Firefox and WebKit run focused core journeys to catch engine-specific behavior without tripling the entire suite.
- Security scenarios run in a separate path-gated, scheduled, and reusable workflow.
- AI teaching changes use deterministic policy tests plus the complete model evaluation gate; focused cases help iteration but cannot establish model eligibility.
- Actual-browser UX audits are mandatory for browser-observable work under the repository harness. Playwright is supporting evidence, not a substitute for experiencing the final flow.
See Development and the agent harness strategy for the working loop.
frontend/ React application, public curriculum, build-time public pages
backend/ Express API, persistence, tutor policy, session/execution control
content/ Backend-only protected learning and memory-warmup material
e2e/ Playwright journeys, fixtures, and security scenarios
swa-api/ Managed crawler share-metadata adapter
supabase/migrations/ Ordered database and RLS source of truth
infra/ Azure Bicep, VM bootstrap, promotion, operations scripts
.github/workflows/ CI, E2E, security, production release, synthetic monitoring
scripts/ Cross-package governance, harness, release, and audit tooling
docs/ Product, architecture, development, quality, and authoring truth
When extending the system, prefer the existing seams: a route module rather than a second server, a registered harness backend rather than inline language branching, a shared tutor renderer rather than per-panel parsing, a forward migration rather than editing history, and an execution-backend implementation rather than leaking infrastructure details into routes.
| Term | Meaning in this system |
|---|---|
| ACI | Azure Container Instances, used only as optional overflow inside the hybrid execution shape |
| BYOK | Bring your own OpenAI key; encrypted server-side and distinct from platform-funded usage |
| Canonical context | Lesson, progress, and teaching state resolved from server-owned sources rather than trusted from the browser |
| Contextual evidence episode | Server-owned identity for one learner, canonical lesson, and normalized run error inside a database-owned claim window; client-selected state cannot mint a fresh boundary, while a genuinely later post-expiry occurrence can qualify again |
| RLS | Postgres row-level security, used as defense in depth for user-scoped data |
| Runner | Ephemeral isolated container that executes one learner session's project |
| SWA | Azure Static Web Apps, which hosts the web client and the crawler-facing share function |
Copyright © 2026 Mehul Srivastava. All rights reserved. See LICENSE.