Skip to content

Guidance for a reproducible ClawBench task-authoring comparison #68

Description

@Shamratha

Context

I am building a reproducible comparison of a native, non-proxy Codex control
against deepseek/deepseek-v4-flash-0731 in two CLI harnesses. The first two
checks create office artifacts; the final check should author and validate a
complete task in the OpenClaw/ClawBench ecosystem. Every arm receives identical
bytes and we record deterministic score, blind quality score, wall time, token
usage, cost, failures, and interventions over at least three matched runs.

The infrastructure is adapted from
https://github.com/lakshitsachdeva/deepseek-trial, but the task/data and final
results will be new.

Proposed Check 3

Have each arm author a complete ClawBench task package from one frozen seed:

  • working title: Background Task Ledger Export Repair
  • target: Tier 3 repo/multi_tool
  • fixture: a small Python task-ledger exporter with independent normalization,
    terminal-state filtering, and completion-deduplication bugs
  • expected agent trajectory: inspect multiple files, implement fixes, run the
    deterministic verifier, and report the result
  • required output: canonical task YAML, asset pack, tests/verifier, and any
    manifest entry required by the current public pipeline
  • scoring: deterministic completion first; an LLM judge must never rescue a
    deterministic failure

This makes the measured work task authoring—not merely replaying a contaminated
public task—and gives all three arms the same frozen OpenClaw/ClawBench source
snapshot.

Guidance requested

Could a maintainer confirm:

  1. the current canonical task/asset-pack schema and directory layout;
  2. the exact validation command we should treat as official;
  3. whether there is an oracle/reference-solution check for newly authored tasks;
  4. the OpenClaw and ClawBench versions/commits that should be pinned together;
  5. whether this proposed seed is useful, or what seed would better exercise an
    end-to-end OpenClaw workflow without duplicating a public Core v1 task; and
  6. whether three runs per arm is adequate for a small study, or whether your
    current methodology requires a different minimum.

I will keep the final source lock ready: false until these details are fixed
and will record the maintainer response and exact commit in the result package.

If an issue is not the right coordination channel, please point me to the
preferred OpenClaw/ClawBench contact.

Metadata

Metadata

Assignees

No one assigned

    Labels

    P3Low-risk cleanup, docs, polish, ergonomics, or speculative feature.clawsweeper:needs-maintainer-reviewClawSweeper marked this issue as needing maintainer review before automation.clawsweeper:needs-product-decisionClawSweeper marked this issue as needing a product or behavior decision.clawsweeper:no-new-fix-prClawSweeper does not recommend queueing a new automated fix PR for this issue.issue-rating: 🌊 off-meta tidepoolIssue quality rating does not apply to this item.

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions