Skip to content

Audit codex-skills inside Codex Lab dogfood #83

Description

@shiny-code-bot

Current Status

State: Complete. PR #496 merged into the upstream candidate as b599ca2c747754bed84d8986b9ff5fb352230b41 on July 29, 2026.

Final candidate evidence:

  • Binding mandatory, delegated, and independent skill-routing guidance is model-independent and guarded under AGENT-1.
  • The restored gpt-5.6-sol skill guidance, cache continuity, project-doc deduplication, and multi-turn scenarios pass in the exact 11-scenario harness.
  • The external-agent lifecycle scenario recovers both end markers from a large file-backed context handoff without inlining or truncation.
  • The exact clean candidate and Every Code 0.6.116 ran the same gpt-5.5/high task, repository snapshot, prompt, tests, and rubric. Both used all three implicit skills, changed no tests, and passed 2/2 tests.
  • Independent Opus review returned PASS/PASS with no blocking finding, dissent, or material quality difference.
  • Strict convergence validation passes with 367 guarded paths, zero violations, zero stale waivers, and all four historical snapshots reproducible.

Machine-readable provenance and paired-result evidence is recorded in the July 29 evidence comment. No confirmed residual gap requires a new follow-up issue.

Next action: none for #83; #428 now owns the remaining candidate freeze and pending-restore decisions.

Blocked by: nothing.
Last verified: July 29, 2026.

Context

We want to audit whether codex-skills guidance works naturally inside Codex Lab once Codex Lab is usable as the active Codex-based coding CLI. The goal is not to copy Every Code behavior back into Codex Lab; it is to find where skills, Codex Lab product affordances, or both need to change so agents have low-friction, correct paths.

Audit Scope After Unblock

  • GitHub helper-first behavior, comments-first issue orientation, PR checks, merge hygiene, and PR babysitting.
  • Local LLM and local-agent flows, including exact prompt control where Codex Lab should avoid unintended harness/system-prompt pollution.
  • Auto Review current-target handling, stale detached findings, bounded detail retrieval, and review-result surfacing.
  • Browser/tool workflows and real-world task recovery.
  • Skill routing, progressive disclosure, and whether skill instructions remain strong enough inside Codex Lab.
  • Whether Codex Lab needs product-side affordances so skills feel like the shortest safe path, not a manual workaround layer.

Acceptance

  • Codex Lab MVP dogfood plan #28 has reached a dogfoodable state.
  • Run real Codex Lab tasks end to end using codex-skills.
  • Capture concrete friction with reproduction notes and decide whether each fix belongs in Codex Lab, codex-skills, or both.
  • Update or split follow-up issues for any confirmed gaps.

Relationships

Acceptance Criteria

  • The exact immutable Codex Lab candidate and Every Code baseline use the same model, reasoning level, repository snapshot, prompts, and acceptance rubrics.
  • Mandatory, delegated, and independent multi-skill routing succeeds without explicit skill names where the contract requires implicit use.
  • GitHub issue planning uses the correct durable issue, native relationship graph, multiline helpers, and status recovery conventions.
  • Multi-turn resume preserves project/skill context and cache stability without duplicate instruction bloat.
  • File-backed bounded context works for large external-agent handoffs without silently truncating required instructions.
  • At least one representative code task produces a passing focused validation result and independently clean review in both harnesses.
  • Codex Lab has no skipped mandatory instructions, lower task pass rate, unrelated edits, or blocking quality findings on the paired subset.
  • Machine-readable artifacts identify the exact binaries, commits, task inputs, outputs, validation, and reviewer disposition.

Metadata

Metadata

Assignees

No one assigned

    Labels

    planDurable planning issueplan:donePlan completed or superseded

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions