[LLM] Add Terminal-Bench evaluation scripts - #23384
mergennachin wants to merge 5 commits into
Conversation
[ghstack-poisoned]
|
Stack from ghstack (oldest at bottom): |
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23384
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit b527ff7 with merge base 0b3d26d ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
[ghstack-poisoned]
Add setup and run scripts for evaluating the LLM server with Harbor and mini-SWE-agent. TOML files under evals/configs select the model and task settings. Harbor handles task downloads, attempts, scores, and trajectories; the launcher starts the server, checks container connectivity, and records token counts and timings. Validated with launcher dry-run, Harbor command validation, Lintrunner, and ShellCheck. Authored with assistance from OpenAI Codex. ghstack-source-id: 3066ca6 ghstack-comment-id: 5964135991 Pull-Request: #23384
[ghstack-poisoned]
Add setup and run scripts for evaluating the LLM server on macOS with Harbor and mini-SWE-agent. TOML files under evals/configs select the model and task settings. Harbor handles task downloads, attempts, scores, and trajectories; the launcher starts the server, checks container connectivity, and records token counts and timings. Validated with launcher dry-run, Harbor command validation, Lintrunner, and ShellCheck. Authored with assistance from OpenAI Codex. ghstack-source-id: e1bb91e ghstack-comment-id: 5964135991 Pull-Request: #23384
[ghstack-poisoned]
Add setup and run scripts for evaluating the LLM server on macOS with Harbor and mini-SWE-agent. TOML files under evals/configs select the model and task settings. Harbor handles task downloads, attempts, scores, and trajectories; the launcher starts the server, checks container connectivity, and records token counts and timings. Validated with launcher dry-run, Harbor command validation, Lintrunner, and ShellCheck. Authored with assistance from OpenAI Codex. ghstack-source-id: fc40a2f ghstack-comment-id: 5964135991 Pull-Request: #23384
[ghstack-poisoned]
Add macOS setup and run scripts for evaluating the LLM server with Harbor and mini-SWE-agent. TOML files under evals/configs select the model, tasks, sampling, and budgets. A generation check from Docker catches incompatible workers before the evaluation starts. The launcher prints trial outcomes and saves token counts and timings, excluding the preflight. Validated setup and a Qwen3-0.6B fix-git smoke run: one completed trial with no infrastructure errors; reward 0.0 because the agent produced no valid tool calls. Verified incompatible-worker rejection before Harbor starts, preflight exclusion from metrics, nondefault temperature configuration, Lintrunner, and ShellCheck. Authored with assistance from OpenAI Codex. ghstack-source-id: 0e9f4c4 ghstack-comment-id: 5964135991 Pull-Request: #23384
|
@claude Review this PR |
|
Claude finished @mergennachin's task in 1m 43s —— View job The scripts look well put together, but there are two real problems: if Detailed reviewPR Review: #23384SummaryThis PR adds macOS setup and run scripts that use Harbor and mini-SWE-agent to run Terminal-Bench 2.0 against the ExecuTorch LLM server. There are two correctness problems in Security
Code Quality
Testing
RecommendationRequest Changes: fix the |
Add setup and run scripts for evaluating the LLM server with Harbor and mini-SWE-agent. TOML files under evals/configs select the model and task settings. Harbor handles task downloads, attempts, scores, and trajectories; the launcher starts the server, checks container connectivity, and records token counts and timings.
Validated with launcher dry-run, Harbor command validation, Lintrunner, and ShellCheck.
Authored with assistance from OpenAI Codex.