-
-
Notifications
You must be signed in to change notification settings - Fork 29
Guidance for a reproducible ClawBench task-authoring comparison #68
Copy link
Copy link
Open
Labels
P3Low-risk cleanup, docs, polish, ergonomics, or speculative feature.Low-risk cleanup, docs, polish, ergonomics, or speculative feature.clawsweeper:needs-maintainer-reviewClawSweeper marked this issue as needing maintainer review before automation.ClawSweeper marked this issue as needing maintainer review before automation.clawsweeper:needs-product-decisionClawSweeper marked this issue as needing a product or behavior decision.ClawSweeper marked this issue as needing a product or behavior decision.clawsweeper:no-new-fix-prClawSweeper does not recommend queueing a new automated fix PR for this issue.ClawSweeper does not recommend queueing a new automated fix PR for this issue.issue-rating: 🌊 off-meta tidepoolIssue quality rating does not apply to this item.Issue quality rating does not apply to this item.
Description
Metadata
Metadata
Assignees
Labels
P3Low-risk cleanup, docs, polish, ergonomics, or speculative feature.Low-risk cleanup, docs, polish, ergonomics, or speculative feature.clawsweeper:needs-maintainer-reviewClawSweeper marked this issue as needing maintainer review before automation.ClawSweeper marked this issue as needing maintainer review before automation.clawsweeper:needs-product-decisionClawSweeper marked this issue as needing a product or behavior decision.ClawSweeper marked this issue as needing a product or behavior decision.clawsweeper:no-new-fix-prClawSweeper does not recommend queueing a new automated fix PR for this issue.ClawSweeper does not recommend queueing a new automated fix PR for this issue.issue-rating: 🌊 off-meta tidepoolIssue quality rating does not apply to this item.Issue quality rating does not apply to this item.
Type
Fields
Priority
None yet
Context
I am building a reproducible comparison of a native, non-proxy Codex control
against
deepseek/deepseek-v4-flash-0731in two CLI harnesses. The first twochecks create office artifacts; the final check should author and validate a
complete task in the OpenClaw/ClawBench ecosystem. Every arm receives identical
bytes and we record deterministic score, blind quality score, wall time, token
usage, cost, failures, and interventions over at least three matched runs.
The infrastructure is adapted from
https://github.com/lakshitsachdeva/deepseek-trial, but the task/data and final
results will be new.
Proposed Check 3
Have each arm author a complete ClawBench task package from one frozen seed:
repo/multi_toolterminal-state filtering, and completion-deduplication bugs
deterministic verifier, and report the result
manifest entry required by the current public pipeline
deterministic failure
This makes the measured work task authoring—not merely replaying a contaminated
public task—and gives all three arms the same frozen OpenClaw/ClawBench source
snapshot.
Guidance requested
Could a maintainer confirm:
end-to-end OpenClaw workflow without duplicating a public Core v1 task; and
current methodology requires a different minimum.
I will keep the final source lock
ready: falseuntil these details are fixedand will record the maintainer response and exact commit in the result package.
If an issue is not the right coordination channel, please point me to the
preferred OpenClaw/ClawBench contact.