Skip to content

Codex Lab MVP dogfood plan #28

Description

@shiny-code-bot

Finish Line

Codex Lab reaches MVP when one provenance-verified current binary supports the three workflows required for real product dogfood:

  1. Automatic Validation runs after relevant code changes, emits bounded structured results, and permits at most one controlled correction cycle. The MVP preserves parser and project-command validation and includes at least one automatically selected changed-file executable provider.
  2. Third-party agents launch through the native agent system with workspace/session provenance, bounded status and output, completion/failure/timeout/cancellation cleanup, and terminal-state recovery proven from the current binary.
  3. Automatic Background Review runs detached after completed code-changing turns, records durable current findings, avoids duplicate/stale noise, and remains distinct from manual Review and Approval Review/Guardian.

History/resume, auth, provenance, TUI status, and upstream substrate work are MVP requirements only where they are necessary to run these workflows safely. Once the three workflows pass deterministic and bounded current-binary smoke gates, Codex Lab begins real dogfood with Every Code retained as rollback. Broad upstream parity, architecture thinning, polish, and the permanent default-route soak are post-MVP unless a dogfood failure promotes a specific item.

Current Status

State: MVP foundations remain complete; this issue stays closed as the durable recovery point for Codex Lab planning.

The supported published candidate remains codex-lab-v0.1.0-lab.4 at source 902f6ddd7a797ce432c6e635552b854dd53ce00a, bundle version 41. Real private dogfood remains active, with private task details retained in their operational repository.

The current release hierarchy is #624#703#707. PR #716 merged on August 17, 2026 at main@780e5f32effbaeacf8a5dccdb0728b5a7bed5d1b; no lab.5 tag or release exists. #707 is the sole active Project Now lane with 16 measurable MultiAgentV2 residual failures. #706 is waiting with its original portability clusters; #708 is Next; #709 is blocked by #708; #710-#713 remain later. #230/#664 stay parked until #624 completes.

Next recovery action: classify the 16 #707 failures from immutable main@780e5f32effbaeacf8a5dccdb0728b5a7bed5d1b with Opus and Gemini and land the smallest coherent product-contract stage. Do not publish lab.5 or activate #230 before #703 records one fully green validation-only release rehearsal.

Source Material

Use restored cbusillo/code issues and commits as historical source material, not as automatic carry-forward status. Important restored planning issues include:

Recent restored commits worth mining include the Codex substrate migration, Codex Desktop startup compatibility, Every Code identity/config alignment, feature inventory, parity ledger, prompt-cache prefix preservation, auto-review proof metrics, auto-review lifecycle/store/ledger work, agent context file preloading, worktree decision gate hardening, and release/build runner fixes.

Base Platform Decision

Codex Lab remains the product base. Build on the Codex CLI/Desktop/app-server substrate because the hard-to-recreate value is Codex Desktop compatibility, Codex iOS/mobile control of Desktop, Codex subscription auth, upstream openai/codex compatibility, and continuity with the Every Code Codex-base port path.

opencode and other coding harnesses are reference implementations, not the base. Steal ideas when they improve Codex Lab without breaking Codex compatibility and can be validated through the exec harness or focused tests.

Guardrails

  • Use fixtures/tests/concepts before broad implementation overlays where feasible.
  • Treat missing Every Code behavior as unclassified until this plan or a child issue records Port, Rewrite, Covered, Defer, or Retire.
  • Keep openai as fetch-only upstream.
  • Avoid direct work on protected/default branches; use focused task branches and PRs for implementation.
  • Every code-bearing port slice should include scoped validation evidence.
  • The exec harness is mandatory for Codex Lab-specific regressions: skills prompt strength, cached-token stability, thread continuity, Desktop/app-server compatibility, and Every Code workflow expectations.
  • Cache-sensitive scenarios should compare input tokens, cached input tokens, cache ratio, and normalized prompt-prefix stability against explicit baselines.

MVP Slices

  1. Repo and runner recovery.

    • Verify cbusillo/codex-lab runners after repo move.
    • Keep cbusillo/code runner-free unless a future archival workflow explicitly needs one.
  2. Planning source of truth.

    • Make this issue the durable parent plan.
    • Create focused child issues for independent MVP slices.
    • Realign repo AGENTS.md so future agents find this issue instead of relying on local plan files.
  3. Exec harness hardening and regression gates.

    • Fix known automation findings: cleanup just argument forwarding and exec-harness workflow path triggers.
    • Add skills-cache-continuity as the first cache-sensitive scenario.
    • Add deterministic fake-Responses coverage first, then local-LLM dogfood and frontier release variants.
  4. Skills and prompt-cache protection.

    • Protect against skills becoming less directive or being treated as optional advice.
    • Protect against prompt-stack churn reducing cached-token reuse.
  5. Auto Review proof loop.

    • Restore actionable review evidence without broad-review noise.
    • Mine restored Every Code auto-review lifecycle/store/ledger commits and issues before implementation.
  6. Agents and third-party agent orchestration.

    • Rebuild configurable agent roles, third-party agents, local LLM roles, and review/validation loops on Codex Lab primitives.
  7. Codex Desktop/app-server compatibility.

    • Validate Codex Lab in Desktop and keep app-server behavior upstream-shaped unless an additive extension is explicitly validated.
  8. Code Bridge/browser and remote-control workflows.

    • Preserve or replace Code Bridge/browser/control capability with clear Desktop/app-server boundaries.
  9. Auto Drive.

    • Rebuild Auto Drive on validated Codex thread/session/worktree/token primitives rather than copying old implementation wholesale.
  10. Local LLM dogfood.

    • Make LM Studio/OpenAI-compatible local endpoints first-class for bounded dogfood basics, then graduate roles only after repeated proof.

Opencode And Other References

opencode is the primary near-term reference for plugin/hooks ergonomics, named agents/subagents, LM Studio flow, provider/auth boundaries, permissions UX, run artifacts, GitHub automation presentation, prompt/context controls, and MCP/tool discovery.

Keep a lightweight watch on OpenHands, SWE-agent/mini-SWE-agent, Aider, Cline/Roo Code, Goose, and Continue for product architecture ideas. Separately, keep Terminal-Bench, promptfoo, Inspect AI, SWE-bench, Aider benchmarks, and BrowserGym/WebArena/OSWorld in mind for eval and grading ideas. These are references only.

Next Actions

  1. Begin bounded opt-in Codex Lab dogfood from merged source 357b8e4ac06053968ba2a5232c4369bcec849458; keep Every Code available as the immediate rollback route.
  2. Record concrete dogfood regressions as focused issues and promote only those defects into blocking work.
  3. Continue Define Every Code-owned convergence contracts #126/Upstream sync and Codex Lab feature reclassification #229 upstream classification and synchronization without retroactively reopening the MVP unless a material workflow regression is found.
  4. Use Cut over to converged Codex Lab baseline and primary harness #343 for paired quality evaluation, rollback/re-entry, multi-day soak, and the eventual permanent default-route decision.
  5. Treat Adopt Background Review terminology across product surfaces #303, Sync TUI, Guardian, and CI substrate #308, Reduce managed approval friction in dogfood workflows #313, Surface structured rate-limit diagnostics for Claude CLI agents #315, and Add a lightweight structured LLM request side channel #319 as post-MVP terminology, presentation, friction, diagnostics, or observability work.

Relationships

Post-Cutover Roadmap

As of July 29, 2026, this closed issue remains the durable MVP evidence and historical recovery point. The active post-cutover roadmap is #498.

Current queue:

Future sessions should recover MVP intent here, then use #498 for current direction and sequencing.

Metadata

Metadata

Assignees

No one assigned

    Labels

    planDurable planning issueplan:donePlan completed or superseded

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions