Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .github/copilot-instructions.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ This is a benchmark for evaluating coding agents on real-world Business Central

- **Dataset**: Benchmark entries following SWE-Bench schema with BC-specific adjustments
- **Python Package** (`src/bcbench/`): CLI tools, agent implementations, and validation utilities
- **Core Library** (`packages/bcbench-core/`): Reusable, distributable evaluation library consumed by `bcbench`; must never import `bcbench`, read environment variables, or contain BC-Bench policy
- **PowerShell Scripts** (`scripts/`): Environment setup and dataset verification using AL-GO/BCContainerHelper
- **Tools** (`tools/`): Ad-hoc scripts for GitHub Artifacts download, etc
- **Agent Evaluations**: Focuses on GitHub Copilot CLI and Claude Code
Expand Down Expand Up @@ -52,7 +53,7 @@ def test_full_metrics_flow_to_success_result(self, sample_context):
```

### Linting and formatting
Ruff is the single source of truth (`uv run ruff check --fix`, `uv run ruff format`); config lives in `pyproject.toml`.
Ruff is the single source of truth (`uv run ruff check --fix`, `uv run ruff format`); the baseline lives in `packages/bcbench-core/pyproject.toml` and the root `pyproject.toml` extends it.
Lean on ruff's default rule set rather than growing `extend-select`, and prefer fixing violations over suppressing them. If a violation is genuinely intentional, use a targeted `# noqa: RULE - rationale` at that line instead of a repo-wide `ignore` entry.

## No Backward compatibility
Expand Down
8 changes: 7 additions & 1 deletion .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,13 @@ repos:
- repo: local
hooks:
- id: ty
name: ty check
name: ty check (bcbench)
entry: uv run ty check
language: system
pass_filenames: false
- id: ty-bcbench-core
name: ty check (bcbench-core)
# `uv check` enables ty's uv integration, which missing-direct-dependency requires
entry: uv check --package bcbench-core --locked --no-sync --preview-features check-command
language: system
pass_filenames: false
15 changes: 9 additions & 6 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,14 +19,17 @@ A very high-level overview of the repository structure:

```
BC-Bench/
├── src/bcbench/ # Evaluation harness — agent orchestration, build/test pipeline, results
├── dataset/ # Benchmark dataset tasks
├── scripts/ # Scripts for container setup & test execution; not needed for local development
├── notebooks/ # Analysis and visualization of results
├── evaluator/ # Braintrust scorer integration, used only when uploading result to Braintrust
└── docs/ # GitHub Page for the leaderboard site
├── src/bcbench/ # Evaluation harness — agent orchestration, build/test pipeline, results
├── packages/bcbench-core/ # Reusable evaluation library consumed by src/bcbench (uv workspace member)
├── dataset/ # Benchmark dataset tasks
├── scripts/ # Scripts for container setup & test execution; not needed for local development
├── notebooks/ # Analysis and visualization of results
├── evaluator/ # Braintrust scorer integration, used only when uploading result to Braintrust
└── docs/ # GitHub Page for the leaderboard site
```

The repository is a [uv workspace](https://docs.astral.sh/uv/concepts/projects/workspaces/) with two Python projects: the `bcbench` application at the root and the `bcbench-core` library. `bcbench` depends on `bcbench-core`; never the reverse. The ruff baseline lives in `packages/bcbench-core/pyproject.toml` and the root config extends it with application-only settings. See [packages/bcbench-core/README.md](packages/bcbench-core/README.md) for the library boundary.

## Setup

Prerequisites:
Expand Down
21 changes: 21 additions & 0 deletions packages/bcbench-core/LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) Microsoft Corporation.

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE
27 changes: 27 additions & 0 deletions packages/bcbench-core/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# bcbench-core

Reusable, strongly typed building blocks for evaluating coding agents on Business Central (AL) tasks. [BC-Bench](https://github.com/microsoft/BC-Bench) is the first consumer; other repositories can build their own evaluation applications on top of it.

> Pre-release: not yet published. The API is being extracted from BC-Bench incrementally and may change without notice.

## Boundary

`bcbench-core` contains only genuinely reusable contracts and operations. It must not contain:

- Benchmark policy: categories, datasets, prompts, scoring thresholds, or workflows
- Repository-specific configuration, integrations, credentials, or internal material
- Reads of global configuration or environment variables; callers pass values explicitly
- Imports of the `bcbench` application

The import and environment rules are enforced by ruff (`banned-api` in [`pyproject.toml`](pyproject.toml)); imports of undeclared dependencies are rejected by ty's `missing-direct-dependency` rule.

## Development

The package is a [uv workspace](https://docs.astral.sh/uv/concepts/projects/workspaces/) member of the BC-Bench repository. From the repository root:

```bash
uv sync --all-groups
uv run ruff check packages/bcbench-core
uv check --package bcbench-core --preview-features check-command
uv build --package bcbench-core
```
84 changes: 84 additions & 0 deletions packages/bcbench-core/pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
[build-system]
requires = ["uv_build>=0.12.19,<0.13"]
build-backend = "uv_build"

[project]
name = "bcbench-core"
version = "0.1.0"
description = "Reusable, strongly typed building blocks for evaluating coding agents on Business Central (AL) tasks"
readme = "README.md"
requires-python = ">=3.13"
license = "MIT"
license-files = ["LICENSE"]
authors = [{ name = "Microsoft Corporation" }]
classifiers = [
# Remove once publishing is decided; PyPI rejects uploads carrying this classifier.
"Private :: Do Not Upload",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3.13",
"Typing :: Typed",
]
dependencies = []

[tool.ruff]
target-version = "py313"
line-length = 200

[tool.ruff.lint]
extend-select = [
# Restore the pyflakes/pycodestyle correctness rules that ruff 0.16 dropped from
# its default set (E711/E712/E713/E714/E721/E731/E741, F403/F405/F406/F722, etc.).
"E4", # pycodestyle: import placement (E401/E402)
"E7", # pycodestyle: statement/comparison lints (== None, lambda assignment, ...)
"F", # pyflakes: undefined names, star-import hazards, unused imports
"PLE", # Pylint errors (includes __all__ validation)
"PLW", # Pylint warnings (includes subprocess.run check requirement)
"I", # isort (import sorting)
"UP", # pyupgrade (modernize Python code)
"B", # flake8-bugbear (find likely bugs)
"SIM", # flake8-simplify (simplify code)
"C4", # flake8-comprehensions (better list/dict comprehensions)
"RET", # flake8-return (simplify return statements)
"PTH", # flake8-use-pathlib (prefer pathlib over os.path)
"RUF", # Ruff-specific rules
"TID", # flake8-tidy-imports (ban relative imports + extensible banned-API list)
"LOG", # flake8-logging: catch deprecated logging.warn, misuse of exception(), etc.
"G", # flake8-logging-format: catch string concat / % formatting in log calls
"TRY", # tryceratops: better exception handling
"PT", # flake8-pytest-style: consistent pytest patterns
"ANN", # flake8-annotations: enforce type hints (we prefer strong typing)
"T20", # flake8-print: prefer logger over print (we have a proper logger)
"FURB", # refurb: modern Python idioms (e.g. x or y over x if x else y)
"ISC", # implicit string concat (catches missing comma in lists of strings)
"ICN", # import conventions (np, pd, plt aliases)
"PYI", # type stubs (no-op now, guard for future)
"SLOT", # require __slots__ on str/tuple/namedtuple subclasses
"ASYNC", # async best practices (no-op now, guard for future)
"PERF", # perflint: avoid needless per-iteration overhead (prefer extend/comprehensions)
"RSE", # flake8-raise: drop redundant parentheses on bare exception raises
"PGH", # pygrep-hooks: require codes on noqa/type-ignore comments (PGH003/PGH004)
"S307", # flake8-bandit: ban eval (was PGH001, which ruff removed in favour of this rule)
"DTZ", # flake8-datetimez: require timezone-aware datetime usage
"NPY", # NumPy-specific correctness and modernization checks
"PIE", # flake8-pie: miscellaneous correctness and simplification checks
]

ignore = [
# Deliberate style choices with many intentional violations; everything else is fixed in-tree
# (targeted `# noqa` with a rationale is preferred over adding entries here).
"G004", # logging-f-string: f-strings are more readable; perf cost is negligible
"TRY003", # raise-vanilla-args: forces a custom exception class for every error message
]

[tool.ruff.lint.flake8-tidy-imports]
ban-relative-imports = "all"

[tool.ruff.lint.flake8-tidy-imports.banned-api]
"bcbench".msg = "bcbench-core must not depend on the BC-Bench application."
"os.environ".msg = "bcbench-core must not read the environment; accept values as parameters."
"os.getenv".msg = "bcbench-core must not read the environment; accept values as parameters."
"dotenv".msg = "bcbench-core must not load configuration; accept values as parameters."

[tool.ty.rules]
# Strictest baseline for the public library; relax a specific rule only with a rationale.
all = "error"
Empty file.
Empty file.
75 changes: 14 additions & 61 deletions pyproject.toml
Original file line number Diff line number Diff line change
@@ -1,19 +1,21 @@
[build-system]
requires = ["setuptools>=61.0"]
build-backend = "setuptools.build_meta"
requires = ["uv_build>=0.12.19,<0.13"]
build-backend = "uv_build"

[project]
name = "bcbench"
version = "0.13.0"
description = "Benchmarking tool for Business Central (AL) ecosystem, inspired by SWE-Bench"
readme = "README.md"
requires-python = ">=3.13,<3.14"
license = {text = "MIT"}
license = "MIT"
license-files = ["LICENSE"]
authors = [
{name = "Microsoft Corporation"}
]
classifiers = ["Private :: Do Not Upload"]
dependencies = [
"bcbench-core",
"jsonschema>=4.0",
"python-dotenv>=1.2.2",
"requests>=2.0",
Expand All @@ -35,17 +37,17 @@ bcbench = "bcbench.cli:app"
# ty's uv integration (TY_UV=1, used for missing-direct-dependency) needs uv >= 0.12.3
required-version = ">=0.12.19"

[tool.uv.workspace]
members = ["packages/bcbench-core"]

[tool.uv.sources]
bcbench-core = { workspace = true }

[[tool.uv.index]]
name = "microsoft-cfs"
url = "https://packagefeedproxy.microsoft.io/pypi/simple"
default = true

[tool.setuptools.packages.find]
where = ["src"]

[tool.setuptools.package-data]
bcbench = ["agent/*.yaml", "agent/pr_review/scripts/*.ps1"]

[tool.pytest.ini_options]
testpaths = ["tests"]
addopts = ["-v", "--strict-markers", "-m", "not e2e"]
Expand All @@ -56,57 +58,8 @@ markers = [
collect_imported_tests = false

[tool.ruff]
target-version = "py313"
line-length = 200

[tool.ruff.lint]
extend-select = [
# Restore the pyflakes/pycodestyle correctness rules that ruff 0.16 dropped from
# its default set (E711/E712/E713/E714/E721/E731/E741, F403/F405/F406/F722, etc.).
"E4", # pycodestyle: import placement (E401/E402)
"E7", # pycodestyle: statement/comparison lints (== None, lambda assignment, ...)
"F", # pyflakes: undefined names, star-import hazards, unused imports
"PLE", # Pylint errors (includes __all__ validation)
"PLW", # Pylint warnings (includes subprocess.run check requirement)
"I", # isort (import sorting)
"UP", # pyupgrade (modernize Python code)
"B", # flake8-bugbear (find likely bugs)
"SIM", # flake8-simplify (simplify code)
"C4", # flake8-comprehensions (better list/dict comprehensions)
"RET", # flake8-return (simplify return statements)
"PTH", # flake8-use-pathlib (prefer pathlib over os.path)
"RUF", # Ruff-specific rules
"TID", # flake8-tidy-imports (ban relative imports + extensible banned-API list)
"LOG", # flake8-logging: catch deprecated logging.warn, misuse of exception(), etc.
"G", # flake8-logging-format: catch string concat / % formatting in log calls
"TRY", # tryceratops: better exception handling
"PT", # flake8-pytest-style: consistent pytest patterns (~30 test files)
"ANN", # flake8-annotations: enforce type hints (we prefer strong typing)
"T20", # flake8-print: prefer logger over print (we have a proper logger)
"FURB", # refurb: modern Python idioms (e.g. x or y over x if x else y)
"ISC", # implicit string concat (catches missing comma in lists of strings)
"ICN", # import conventions (np, pd, plt aliases)
"PYI", # type stubs (no-op now, guard for future)
"SLOT", # require __slots__ on str/tuple/namedtuple subclasses
"ASYNC", # async best practices (no-op now, guard for future)
"PERF", # perflint: avoid needless per-iteration overhead (prefer extend/comprehensions)
"RSE", # flake8-raise: drop redundant parentheses on bare exception raises
"PGH", # pygrep-hooks: require codes on noqa/type-ignore comments (PGH003/PGH004)
"S307", # flake8-bandit: ban eval (was PGH001, which ruff removed in favour of this rule)
"DTZ", # flake8-datetimez: require timezone-aware datetime usage
"NPY", # NumPy-specific correctness and modernization checks
"PIE", # flake8-pie: miscellaneous correctness and simplification checks
]

ignore = [
# Deliberate style choices with many intentional violations; everything else is fixed in-tree
# (targeted `# noqa` with a rationale is preferred over adding entries here).
"G004", # logging-f-string: f-strings are more readable; perf cost is negligible (143 sites)
"TRY003", # raise-vanilla-args: forces a custom exception class for every error message (96 sites)
]

[tool.ruff.lint.flake8-tidy-imports]
ban-relative-imports = "all"
# The app inherits the bcbench-core lint baseline; tables redefined below replace the inherited ones.
extend = "packages/bcbench-core/pyproject.toml"

[tool.ruff.lint.flake8-tidy-imports.banned-api]
# Force sandboxed rendering:
Expand All @@ -120,7 +73,7 @@ ban-relative-imports = "all"
"tools/**" = ["T20"] # standalone CLI scripts

[tool.ty.src]
exclude = ["notebooks/"]
exclude = ["notebooks/", "packages/"]

[tool.ty.analysis]
# bc-eval[capi] is internal and installed only into a separate venv by bcal-evaluation.yml
Expand Down
13 changes: 13 additions & 0 deletions uv.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading