Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions examples/llm_server/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ examples/llm_server/
spec/ # language-neutral OpenAI contract ExecuTorch targets
conformance/ # one test suite every language server must pass
python/ # Python server implementation (current)
evals/ # task accuracy and performance through the server
# cpp/ # future: no-Python single-binary server
```

Expand Down Expand Up @@ -106,3 +107,8 @@ Reliability guidance:
`tools` were included in the request.
- If a request fails with `unsupported_parameter`, remove or disable that
OpenAI knob in your pi/client config.

## Evaluate with Terminal-Bench

See [Terminal-Bench](evals/terminal_bench/README.md) for setup, configuration,
and evaluation commands on macOS.
5 changes: 5 additions & 0 deletions examples/llm_server/evals/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# Copyright (c) Meta Platforms, Inc. and affiliates.
# All rights reserved.
#
# This source code is licensed under the BSD-style license found in the
# LICENSE file in the root directory of this source tree.
1 change: 1 addition & 0 deletions examples/llm_server/evals/configs/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
*.local.toml
17 changes: 17 additions & 0 deletions examples/llm_server/evals/configs/terminal-bench.example.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Copy to terminal-bench.local.toml and set your model paths.
[server]
python = "/path/to/executorch-env/bin/python"
worker_bin = "/path/to/model_worker"
model_path = "/path/to/model.pte"
tokenizer_path = "/path/to/tokenizer.json"
hf_tokenizer = "/path/to/pinned-hf-tokenizer"
model_id = "qwen3"
max_context = 8192
no_think = true

[terminal_bench]
tasks = ["fix-git"]
max_output_tokens = 1024
temperature = 0.0
step_limit = 100
attempts = 1
25 changes: 25 additions & 0 deletions examples/llm_server/evals/terminal_bench/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
# Terminal-Bench (macOS)

Prepare a worker, exported model, and [server environment](../../python/README.md).
Use a worker built for the server checkout; worker protocols can differ between revisions.
From the repository root:

```bash
cd examples/llm_server/evals/terminal_bench
bash setup.sh
cp ../configs/terminal-bench.example.toml ../configs/terminal-bench.local.toml
# Edit the model paths and context limit in the local TOML.
bash run.sh --config ../configs/terminal-bench.local.toml
```

Setup installs Colima (requires Homebrew) and Harbor, reusing Docker if available.
Harbor downloads Terminal-Bench 2.0 tasks and runs mini-SWE-agent in containers.

The TOML selects the model, tasks, attempts, temperature, and token budgets.
`max_context` must fit the exported model. The default `fix-git` task is a smoke
check; use a model capable of tool calling for meaningful scores.

Scores, trajectories, logs, and token/timing metrics are saved under
`~/.cache/executorch-evals/terminal-bench/runs/`. Use `--output DIR` to choose a
results directory or `--dry-run` to inspect commands before running.
Metrics exclude the generation preflight.
5 changes: 5 additions & 0 deletions examples/llm_server/evals/terminal_bench/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# Copyright (c) Meta Platforms, Inc. and affiliates.
# All rights reserved.
#
# This source code is licensed under the BSD-style license found in the
# LICENSE file in the root directory of this source tree.
2 changes: 2 additions & 0 deletions examples/llm_server/evals/terminal_bench/requirements.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
# Harbor runs separately from the model server and requires Python 3.12+.
harbor==0.22.0
19 changes: 19 additions & 0 deletions examples/llm_server/evals/terminal_bench/run.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
#!/usr/bin/env bash
# Copyright (c) Meta Platforms, Inc. and affiliates.
# All rights reserved.
#
# This source code is licensed under the BSD-style license found in the
# LICENSE file in the root directory of this source tree.
set -euo pipefail

evals_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)
repo_dir=$(cd -- "${evals_dir}/../../.." && pwd)
export EXECUTORCH_EVAL_CACHE="${EXECUTORCH_EVAL_CACHE:-${HOME}/.cache/executorch-evals}"

eval_python="${EXECUTORCH_EVAL_CACHE}/terminal-bench/venv/bin/python"
if [[ ! -x "${eval_python}" ]]; then
echo "Run bash ${evals_dir}/terminal_bench/setup.sh first." >&2
exit 2
fi
export PYTHONPATH="${repo_dir}/src${PYTHONPATH:+:${PYTHONPATH}}"
exec "${eval_python}" -m "executorch.examples.llm_server.evals.terminal_bench.runner" "$@"
Loading
Loading