diff --git a/docs/CONTEXTUAL_LEARNING_AND_DELIVERY_VELOCITY_PLAN.md b/docs/CONTEXTUAL_LEARNING_AND_DELIVERY_VELOCITY_PLAN.md index ed67003..d80455a 100644 --- a/docs/CONTEXTUAL_LEARNING_AND_DELIVERY_VELOCITY_PLAN.md +++ b/docs/CONTEXTUAL_LEARNING_AND_DELIVERY_VELOCITY_PLAN.md @@ -1,6 +1,6 @@ # Contextual Learning and Delivery Velocity Plan -> **Status:** Every previously eligible engineering lane is implemented and green through B8. Release 1C entered engineering validation on 2026-08-30 under a dated founder exception that waives waiting for real-learner evidence but preserves its security, cost, quality, accessibility, CI, review, rollback, and production-browser gates. Human judge calibration and learner outcomes remain explicit non-claims. B6 still requires its DAU trigger. +> **Status:** Every previously eligible engineering lane is implemented and green through B8. Release 1C merged, deployed, and passed exact-SHA production verification on 2026-08-31 under the dated founder exception that waived waiting for real-learner evidence while preserving its security, cost, quality, accessibility, CI, review, rollback, and production-browser gates. Human judge calibration and learner outcomes remain explicit non-claims. B6 still requires its DAU trigger. > > **Prepared:** 2026-07-30 > @@ -434,6 +434,18 @@ not promote production. The live rollback/forward-promotion drill is a separate post-merge operational obligation required before declaring the production gate fully exercised. +**Live-drill completion (2026-08-31):** rollback run +[`33409584774`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33409584774) +promoted retained candidate `297a09b4d4873d10e1b1e688f151f8a3a325fcb3` +and passed independent frontend-identity and backend deep-health checks in 2m39s. +Forward release +[`33409904066`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33409904066) +restored exact current candidate `4edcd73fa3a8669ac09495fbb7d5234d8ccfc742`; +production synthetic +[`33411478984`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33411478984) +and in-app Browser phase audit `fcaf8da5-d989-46be-8f94-3f634c469050` +then passed. No database migration existed between the two candidates. + Deliver: - build immutable frontend and backend artifacts once; @@ -699,15 +711,33 @@ Exit: ### Release 1C — Contextual tutor offer -**Founder exception and implementation status (2026-08-30):** engineering work -is in final validation on `dev/contextual-tutor-1c`. The founder explicitly -waived waiting for real-learner experiment evidence. The exception does not -weaken the security, cost, AI-quality, deterministic, accessibility, -live-browser, CI, review, rollback, or production-verification gates. The -two-human judge calibration is not complete and is not claimed. See -`docs/RELEASE_1C_ENTRY_GATE.md` for the decision record and current matrix. - -**Entry gate:** Release 1B's preregistered experiment passes the primary recovery rule and every applicable guardrail in Section 10.3; B2, eval v2, authority, idempotency, cost, and security gates also pass. Five qualitative sessions alone cannot unlock 1C. +**Founder exception and delivery status (2026-08-31):** Release 1C merged in +PR #41, its evidence-scoped recovery follow-up merged in PR #42, and both exact +SHAs passed release, production synthetic, and interactive production-browser +verification. The founder explicitly waived waiting for real-learner +experiment evidence. The exception did not weaken the security, cost, +AI-quality, deterministic, accessibility, live-browser, CI, review, rollback, +or production-verification requirements. Automated rollback contracts passed, +and Release 0P's separate live production rollback/forward-promotion drill +subsequently passed on 2026-08-31 with exact identity, health, synthetic, and +Browser evidence. The two-human judge calibration remains incomplete and is +not claimed. See +`docs/RELEASE_1C_ENTRY_GATE.md` for the decision record and evidence matrix. + +**Applied entry gate:** the dated founder exception superseded Release 1B's +preregistered learner experiment and five-session prerequisites for this +engineering release. B2, deterministic AI-safety validators, the complete +automated eval-v2 run, authority, idempotency, cost, security, accessibility, +CI, review, automated rollback controls, and production journey verification +remained mandatory and passed. The separate live 0P rollback/forward-promotion +drill also passed on 2026-08-31, restoring the exact current candidate before +synthetic and live-browser acceptance. The automated model run was provisional risk evidence, not +authoritative eval-v2 approval: Section 9.2's separate two-human judge +calibration was not performed and remains required before automated judging can +support an authoritative quality claim. The engineering release was supported +by deterministic safeguards, provisional automated evidence, an independent +kill switch, exact-SHA deployment, and production-browser verification; it does +not establish calibrated human agreement, learner recovery, or retention. **Complexity:** two to four weeks. @@ -1171,7 +1201,10 @@ risk decision. Release 0D is now implemented on this branch, so server authority, atomic admission, answer-leak protection, output safety, and model eligibility are non-bypassable. The production disable/rollback drill target remains under 10 minutes and must be recorded as operational evidence rather -than inferred from automated tests. +than inferred from automated tests. The 2026-08-31 live drill completed +rollback in 2m39s, forward-promoted the exact current candidate, and passed +synthetic plus Browser acceptance; its evidence is recorded in the Release 0P +section and `docs/RELEASE_1C_ENTRY_GATE.md`. ## 10. Success and stop gates @@ -1243,13 +1276,13 @@ list of unfinished engineering work. ### P0 -1. **0P — implementation complete:** approved candidate manifest and exact artifact promotion; the post-merge live rollback/forward drill remains operational evidence. +1. **0P — complete:** approved candidate manifest and exact artifact promotion; the 2026-08-31 live rollback/forward drill passed with recorded identity, health, synthetic, and Browser evidence. 2. **0B — complete:** suffix-scoped E2E teardown, non-test-user deletion guard, and overlapping-run proof. 3. **1A — complete:** late Run/Check/tutor operation identity and stale-result rejection. 4. **0D — complete:** atomic platform-AI reservation/admission and complete eval gate v2. 5. **0C — complete:** first-run complete-answer rescue removed. 6. **0A — engineering complete:** internal non-counting share preview path and truthful share outcomes; real-destination production unfurls remain external evidence. -7. **0D before 1C — prerequisite complete:** server-authoritative authenticated lesson/mastery context; 1C remains held by its other conjunctive gates. +7. **0D before 1C — prerequisite and delivery complete:** server-authoritative authenticated lesson/mastery context shipped before Release 1C; the later founder exception and verified 1C delivery are recorded in `RELEASE_1C_ENTRY_GATE.md`. ### P1 @@ -1261,7 +1294,7 @@ list of unfinished engineering work. ### P2 12. **B1/B2 — complete for engineering:** locked memory and Socratic groundwork; real-user outcomes remain pending. -13. **1C — correctly held:** contextual tutor offer cannot start until every entry gate passes. +13. **1C — engineering release complete:** the founder exception waived the learner-evidence prerequisites for this release; automated quality, security, cost, accessibility, CI, review, rollback-contract, deployment, production-journey, and live 0P rollback/forward-promotion gates passed. Human calibration and learner outcomes remain explicit non-claims. 14. **B3/B7/B8 — engineering complete:** real-user dropoff and other product outcomes remain pending. 15. **Reporting — engineering controls complete:** real-traffic cost/performance/outcome reporting begins only when traffic exists. 16. **Visual polish — completed for the eligible surfaces in this workstream:** future polish remains ordinary product-roadmap work. diff --git a/docs/CONTEXTUAL_LEARNING_ROADMAP_FINAL_AUDIT.md b/docs/CONTEXTUAL_LEARNING_ROADMAP_FINAL_AUDIT.md index 08ce66a..4a31da0 100644 --- a/docs/CONTEXTUAL_LEARNING_ROADMAP_FINAL_AUDIT.md +++ b/docs/CONTEXTUAL_LEARNING_ROADMAP_FINAL_AUDIT.md @@ -1,5 +1,11 @@ # Contextual learning roadmap — final engineering audit +> **Historical cutoff notice (updated 2026-08-31):** This document records the +> 2026-07-31 gate state. Its Release 1C hold was later superseded by the dated +> founder exception and the verified delivery recorded in +> [`RELEASE_1C_ENTRY_GATE.md`](RELEASE_1C_ENTRY_GATE.md). Do not use the 1C row +> below as the current release status. + Status: every currently eligible engineering lane is complete; deliberately gated learner, traffic, post-merge observation, and future-phase work remains held rather than being claimed without evidence @@ -25,8 +31,10 @@ the full retained browser suite have current evidence. Five categories are intentionally not represented as completed product or operational outcomes: -1. Release 1C cannot start until its locked learner experiment, two-human eval - calibration, named ownership, and approvals exist. +1. At this audit's cutoff, Release 1C could not start until its locked learner + experiment, two-human eval calibration, named ownership, and approvals + existed. The later founder exception and delivery record supersede this + historical hold without claiming human calibration or learner outcomes. 2. B6 cannot ship until seven-day-average DAU reaches 100. No dated metric artifact proving that trigger exists, so the gate remains closed; if it does not fire, the locked roadmap carries B6 into Phase C unchanged. @@ -52,7 +60,7 @@ Cinematic duration remains paused exactly as requested. | **0D — AI trust** | Complete | `7d4b4fb` plus provenance/test fixes provides atomic admission, server-owned trusted context, evaluated-model enforcement, output safety, and authoritative eval gates. Required migrations are applied to the linked development Supabase project; production migration remains part of promotion. | | **1A — context correctness** | Complete | `b03dc13` rejects stale Run, Check, tutor, selection, completion, and stdin results by revision and operation identity without adding an AI request. | | **1B — deterministic guide** | Engineering/internal-dogfood complete | `cf9f0a7` supplies the default-off, authored repeated-error guide with one attention owner and no automatic AI call. Browser, phone-keyboard, accessibility, and reduced-motion evidence pass. The five-session learner gate remains held, so external rollout is not claimed. | -| **1C — contextual tutor offer** | Correctly not started | `e811e20` and `RELEASE_1C_ENTRY_GATE.md` prove the conjunctive entry gate is not met. Implementing it now would violate the approved plan. | +| **1C — contextual tutor offer** | Correctly not started at this audit's cutoff; superseded | `e811e20` captured the 2026-07-31 hold. The current status and later founder exception are recorded in `RELEASE_1C_ENTRY_GATE.md`. | | **1D — CI shadow pilot** | Additive implementation complete | `ded15ae` through `0398a4e` adds risk metadata, the advisory critical lane, frozen catch corpus, retained full browser coverage, and measured six-shard execution. The post-merge observation window has not begun and no demotion decision has been made. | | **B1 — memory read-side** | Complete | `c81c514` through `2a0f6b9` adds server-scored retrieval evidence, own-user RLS, backend-only writes, export/deletion, authored warm-ups, and failure recovery. It does not claim real-user D7 improvement or expose the Phase C mastery graph. | | **B2 — Socratic default** | Complete | `9994ae4` through `74c99b6` enforces one clarifying question first, bounded later help, and no complete answer across scripted/model and auth/anonymous paths; the 60-case gate, preview, persona, and remote checks pass. | diff --git a/docs/RELEASE_0P_RUNBOOK.md b/docs/RELEASE_0P_RUNBOOK.md index 4d0c9b6..98b6756 100644 --- a/docs/RELEASE_0P_RUNBOOK.md +++ b/docs/RELEASE_0P_RUNBOOK.md @@ -190,3 +190,26 @@ previous successful 0P candidate and then forward-promote the latest candidate again. Record both run URLs in the release issue/audit. The automated failure harness is required pre-merge; the live drill is required before declaring the operational release gate fully exercised. + +## Live drill record + +The required first controlled production drill passed on 2026-08-31: + +- rollback run + [`33409584774`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33409584774) + promoted successful retained candidate + `297a09b4d4873d10e1b1e688f151f8a3a325fcb3` in 2m39s; +- cache-busted frontend identity and backend deep health independently passed on + the rolled-back candidate; +- forced forward release + [`33409904066`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33409904066) + restored exact current candidate + `4edcd73fa3a8669ac09495fbb7d5234d8ccfc742` after the complete CI, E2E, and + security gate; +- production synthetic + [`33411478984`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33411478984) + passed on the restored SHA; +- in-app Browser phase audit `fcaf8da5-d989-46be-8f94-3f634c469050` + passed the anonymous Contextual Tutor journey at desktop and 390×844, + duplicate-click admission, focus restoration, and zero-console-error checks; +- no database migration existed between the rollback and forward candidates. diff --git a/docs/RELEASE_1C_ENTRY_GATE.md b/docs/RELEASE_1C_ENTRY_GATE.md index 122fe49..99b89ad 100644 --- a/docs/RELEASE_1C_ENTRY_GATE.md +++ b/docs/RELEASE_1C_ENTRY_GATE.md @@ -1,23 +1,30 @@ # Release 1C contextual tutor decision record -Status: **LOCALLY VERIFIED UNDER FOUNDER EXCEPTION; PR AND DEPLOYMENT PENDING** +Status: **MERGED, DEPLOYED, PRODUCTION-VERIFIED, AND LIVE-ROLLBACK-VERIFIED +UNDER FOUNDER EXCEPTION** Decision date: 2026-08-30 -Branch: `dev/contextual-tutor-1c` +Delivery: [PR #41](https://github.com/msrivas-7/CodeTutor-AI/pull/41), +[recovery follow-up #42](https://github.com/msrivas-7/CodeTutor-AI/pull/42) ## Decision The founder explicitly directed the team to implement Release 1C without waiting for real-learner experiment evidence. That dated exception waives the powered Release 1B learner experiment and five-session rollout prerequisites -for this engineering slice. It does not weaken the security, cost, AI-quality, -accessibility, deterministic-test, live-browser, CI, review, rollback, or -production-verification requirements. +for this engineering slice. Security, cost, deterministic AI safety, the +complete automated model run, accessibility, live-browser, CI, review, +automated rollback controls, and production-verification evidence remained +mandatory. Release 0P's separate live production rollback/forward-promotion +drill subsequently passed on 2026-08-31 and restored the exact current +candidate before synthetic and Browser acceptance. The unperformed Section 9.2 +human calibration was not silently treated as passed: the automated judge +remains provisional rather than authoritative. -The implementation may merge and deploy only when its complete automated model +The implementation merged and deployed only after its complete automated model gate, deterministic suites, adversarial browser journey, PR checks, and review -threads are green. The independent two-human calibration described in Section +threads were green. The independent two-human calibration described in Section 9.2 has not been performed; therefore this release makes no claim that model judging is a calibrated substitute for human evaluation and no learner-outcome, retention, or market-impact claim. @@ -26,15 +33,19 @@ retention, or market-impact claim. | Requirement | Status | Evidence or remaining proof | | --- | --- | --- | -| Founder exception for learner evidence | **Recorded** | 2026-08-30 direction: implement 1C and do not wait for real-user evidence. | +| Founder exception for learner evidence | **Recorded** | The dated 2026-08-30 founder direction is recorded in this decision's `Decision` section and was limited to the Release 1B learner experiment and five-session prerequisites. | | Locked B2 teaching contract | **Retained** | Contextual turns still pass the complete-answer firewall and provide one bounded question/hint. | -| Server authority and stale evidence | **Verified locally** | The server reconstructs and validates the authored move, lesson, revision, evidence code/path/line, and scaffold level. Python is parsed before execution, and only the server-owned compile diagnostic qualifies; stale offers and forged runtime stderr are rejected by the full backend suite. | -| Explicit consent and bounded admission | **Verified locally** | No AI request occurs before `Help me spot it`; one accepted evidence episode can schedule at most one request. The replay identity uses only server-verified actor, canonical lesson, and normalized server error; client epoch/revision cannot reset it. A real-browser double-click produced one turn and one quota decrement, while real Postgres rejected simultaneous disjoint receipt subsets and allowed fresh help only after the database-owned 15-minute window expired. | -| Deterministic recovery | **Verified locally** | Loading, unavailable, kill-switch, stale-generation, and changed-code paths remain useful. The actual browser discarded a stale answer and restored the quota without placing stale content in the transcript. | -| Contextual AI quality | **Passed** | Six contextual golden/adversarial cases cover normal help, source injection, stderr injection, answer pressure, stale history, and line accuracy. Artifact `2026-08-31T04-51-10-727Z-v2.json` passed 72/72 with every intent at 100%, zero deterministic failures, and contract `c0866e3b97a0a7e8…`. | +| Server authority and stale evidence | **Verified** | The server reconstructs and validates the authored move, lesson, revision, evidence code/path/line, and scaffold level. Python is parsed before execution, and only the server-owned compile diagnostic qualifies; stale offers and forged runtime stderr are rejected by the full backend suite. | +| Explicit consent and bounded admission | **Verified** | No AI request occurs before `Help me spot it`; one accepted evidence episode can schedule at most one request. The replay identity uses only server-verified actor, canonical lesson, and normalized server error; client epoch/revision cannot reset it. A real-browser double-click produced one turn and one quota decrement, while real Postgres rejected simultaneous disjoint receipt subsets and allowed fresh help only after the database-owned 15-minute window expired. | +| Deterministic recovery | **Verified** | Loading, unavailable, kill-switch, stale-generation, and changed-code paths remain useful. The actual browser discarded a stale answer and restored the quota without placing stale content in the transcript. PR #42 additionally verified safe consumed-episode recovery and distinct-error re-arming. | +| Contextual AI quality | **Provisional automated evidence passed** | Six contextual golden/adversarial cases cover normal help, source injection, stderr injection, answer pressure, stale history, and line accuracy. Local artifact `backend/eval/runs/2026-08-31T04-51-10-727Z-v2.json` passed 72/72 with every intent at 100%, zero deterministic failures, and contract `c0866e3b97a0a7e8…`. Without Section 9.2 human calibration, this is release-risk evidence rather than authoritative eval-v2 approval. | | Cost | **Within provisional guardrail** | Focused live evidence measured about $0.0077 per accepted call and roughly 280 net-new input tokens; the client and server cap the episode at one call. | | Kill switch | **Implemented** | `contextual_tutor_enabled` is independently operator-controlled and defaults safely when unavailable. | -| Actual browser UX | **Passed locally** | Desktop and 390×500 light/dark/reduced-motion evidence proves explicit consent, one-call double-click protection, current receipt/citation, useful bounded help, retained focus, dismissal persistence, stale recovery, and simultaneous cue/target visibility. The release-boundary replay also proves forged runtime stderr cannot show or fund contextual help while genuine parser diagnostics still can. Finding audit `8a292b5d-dcb6-4678-9e24-11c330d0d987`; phase audit `41617175-7bc9-4905-982c-df5cf735438e`. | +| Actual browser UX | **Passed locally and in production** | Local desktop and 390×500 light/dark/reduced-motion evidence proves explicit consent, one-call double-click protection, current receipt/citation, useful bounded help, retained focus, dismissal persistence, stale recovery, and simultaneous cue/target visibility. Finding audit `8a292b5d-dcb6-4678-9e24-11c330d0d987`; phase audit `41617175-7bc9-4905-982c-df5cf735438e`. The deployed anonymous journey then passed the changed-error offer, grounded response, replay refusal, private recovery, and distinct-error re-arm at desktop and 390×844. | +| PR, review, and CI | **Passed** | PRs #41 and #42 merged with every check green and zero unresolved actionable review threads. | +| Deployment and production | **Passed** | Exact releases [`33363783481`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33363783481) and [`33370195367`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33370195367) succeeded. Exact-SHA production synthetics [`33365791874`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33365791874) and [`33370947310`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33370947310) passed. | +| Rollback controls and rehearsal | **Passed** | Release and rollback workflow contracts passed in CI. Live rollback [`33409584774`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33409584774) promoted retained `297a09b4d4873d10e1b1e688f151f8a3a325fcb3` and completed in 2m39s; independent frontend identity and backend deep health passed. Forward release [`33409904066`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33409904066) restored exact current `4edcd73fa3a8669ac09495fbb7d5234d8ccfc742`; synthetic [`33411478984`](https://github.com/msrivas-7/CodeTutor-AI/actions/runs/33411478984) and in-app Browser phase audit `fcaf8da5-d989-46be-8f94-3f634c469050` then passed at desktop and 390×844 with zero console errors. | +| Observed exposure | **Enabled on verification date** | A fresh anonymous production trial on 2026-08-31 displayed `Help me spot it` after two changed genuine parser failures and returned one structured, line-grounded hint after explicit consent. This dated observation is not a promise that the operator kill switch or rollout scope will never change. | | Human judge calibration | **Not performed; explicitly not claimed** | Two independent human reviewers have not labeled the required stratified set. This is recorded as a non-claim rather than fabricated evidence. | ## Non-negotiable release contract @@ -58,5 +69,5 @@ retention, or market-impact claim. - This exception is not evidence of learner recovery or retention lift. - Persona review and model judges are not real users or two-human calibration. - Passing focused cases cannot replace the complete unfiltered model gate. -- Deployment cannot be called complete until the exact deployed SHA and live - production journey are verified. +- Production verification proves engineering delivery, not learner recovery, + retention, or market impact.