An open-source, forkable engine that runs a tool-enabled GitHub Copilot CLI review over the diff of a Business Central (AL) pull request and posts structured findings as inline PR comments.
The engine is mechanism only. All review knowledge — the skills that
decide what to look for and how to report it — lives in
microsoft/BCQuality. Each consuming
repository owns its policy (which BCQuality repo/ref/layers/skills to use,
severity thresholds) via a small bcquality.config.yaml.
Consumer repo (policy) ──uses──▶ this engine (mechanism) ──clone+filter──▶ BCQuality (knowledge)
| Path | Purpose |
|---|---|
agents/ALReviewAgent/scripts/Invoke-CopilotPRReview.ps1 |
Orchestrator: checks out the PR head, builds the BCQuality task-context, runs the Copilot CLI, parses the findings, renders and posts inline comments + a summary. |
agents/ALReviewAgent/scripts/Get-BCQualityConfig.ps1 |
Loads bcquality.config.yaml and applies environment-variable overrides. |
agents/ALReviewAgent/scripts/Invoke-BCQualityFilter.ps1 |
Prunes a BCQuality clone on disk per the resolved allow/deny/layers policy. |
agents/ALReviewAgent/bcquality.config.yaml |
Default policy baseline. Consumers point at their own copy instead. |
agents/ALReviewAgent/copilot-cli-compatibility.psd1 |
Exact Copilot CLI releases the engine has compatibility-validated, including their OTel CLI-version behavior. |
.github/workflows/review.yml |
Reusable (workflow_call) workflow that wires the whole thing together. |
Online Evals/ |
Pull-based scoring pipeline for evaluating review quality. |
Add a thin caller workflow to your repository. The recommended pattern uses an
unprivileged pull_request intake workflow that saves PR metadata, then a
workflow_run workflow that calls this reusable workflow on the trusted base:
# .github/workflows/pr-review-runner.yml
name: PR Review Runner
on:
workflow_run:
workflows: [PR Review Intake]
types: [completed]
permissions:
contents: read
pull-requests: write
issues: write
copilot-requests: write
jobs:
review:
if: github.event.workflow_run.conclusion == 'success'
uses: microsoft/BC-ALReviewAgent/.github/workflows/review.yml@<pinned-sha>
with:
target_repo: ${{ github.repository }}
engine_ref: <pinned-sha> # keep in sync with the uses: SHA
config_path: .github/bcquality.config.yamlThe engine resolves the PR coordinates from the caller's workflow_run payload
automatically. To review a specific PR (e.g. from workflow_dispatch), pass
pr_number, head_sha, and base_ref explicitly to bypass resolution.
Engine tags use X.Y.Z. X.Y comes from the repo-root VERSION
file and identifies the orchestrator contract/implementation; Z is the
monotonically increasing BCQuality content minor from
bcquality.version in agents/ALReviewAgent/bcquality.config.yaml.
After a merge to main, .github/workflows/version.yml creates an immutable
X.Y.Z tag only when that version does not already exist, then force-moves the
floating latest tag to the merge commit. A merge that changes neither
VERSION nor bcquality.version does not publish a new tag or move latest.
Consumers may pin an immutable tag for reproducibility or follow @latest.
Keep the reusable workflow ref and its engine_ref input aligned so both jobs
check out the same engine version.
| Input | Default | Meaning |
|---|---|---|
target_repo |
caller repo | owner/repo to review and comment on. |
engine_ref |
main |
Ref of this engine repo to check out for scripts. Pin it. |
config_path |
(empty) | Path to the consumer's bcquality.config.yaml, relative to the target repo root. Empty uses the engine default. |
pr_number / head_sha / base_ref |
(empty) | Explicit PR coordinates; bypasses workflow_run resolution. |
minimum_severity |
Medium |
Lowest severity to report (Critical/High/Medium/Low). |
copilot_model |
gpt-5.6-sol |
Root model for final self-review and consolidation; consumers may override it. |
copilot_leaf_model |
gpt-5.6-luna |
Model for isolated review leaves; consumers may override it. |
BCQuality policy (bcquality_repo, bcquality_ref, enabled_layers,
disabled_skills, knowledge_allow, knowledge_deny) and reviewer behaviour
(copilot_model, max_findings_per_domain, fail_on_parse_error, …) can also
be overridden per-input. See .github/workflows/review.yml for the full list.
BCQuality owns each finding's human-readable domain label. The orchestrator
prefers a non-empty findings[].domain value (and accepts PowerShell's
capitalized Domain spelling), then renders and groups that label without
maintaining a duplicate domain taxonomy. For compatibility with older
BCQuality refs, findings without an emitted label fall back to the legacy
from-sub-skill/from_sub_skill map in Invoke-CopilotPRReview.ps1, and
unknown sub-skills fall back to Other. Agent findings retain an explicitly
emitted domain; only unlabeled legacy agent findings use the Agent fallback.
The shared DO schema keeps domain optional for legacy producers, but current
BCQuality review leaves emit a trimmed, single-line, control-free, non-empty
short display label on every finding, including leaf agent findings.
al-code-review preserves that optional value verbatim during rollup and does
not derive or overwrite it from from-sub-skill; its own cross-cutting findings
use exactly Agent.
Accordingly, this consumer never replaces a present non-empty label. It consults
the legacy map, then Other, only when the producer label is absent or empty.
After outer whitespace is trimmed, domain identity is otherwise lossless:
internal whitespace, punctuation, case, and Unicode representation remain
significant. New comment metadata encodes those exact UTF-8 bytes; only legacy
single-token metadata and headings use a separate lowercase compatibility path.
The reusable workflow accepts three levels of BCQuality configuration:
Deterministic leaf execution requires BCQuality commit
b74967bc5b7a454eae19d6a1250199afd869f064 or a newer ref. This is the merge
commit for BCQuality#182,
which introduced the findings-report and skill-index schemas used by the
orchestrator. BCQuality ships one findings-report schema for both roles, so
sub-results and skipped-sub-skills are optional there and nested
sub-results recurse into the whole report. Before running leaves, the
orchestrator derives two role contracts from that pinned schema, writes them as
_review-findings-report.leaf.schema.json and
_review-findings-report.root.schema.json, and points each process at its own
contract:
- The leaf contract removes the super-skill-only
sub-resultsandskipped-sub-skillsproperties. The sharedadditionalProperties: falsethen rejects them, even as empty arrays. - The root contract requires
sub-resultsand validates each entry against the embedded leaf contract instead of recursing into the root contract.
Both are mechanical transforms of the pinned schema, not copies. If a BCQuality
schema change removes the shape the transforms depend on, the run fails closed.
The engine validates reports against its in-memory copies, not the files that
review processes can write. It never strips or rewrites role-violating fields,
and a separate role assertion still names any super-skill-only field found in a
leaf or nested sub-result. A leaf that produces malformed or schema-invalid JSON is recorded
as failed and the remaining leaves continue; model-substitution and telemetry
integrity failures remain fail-closed. _run-manifest.json records leaf and
consolidated-report coverage with a top-level partial status when any usable
review is incomplete; its existing per-process records identify failed leaf IDs
and reasons. If every leaf fails,
the run stops before root consolidation and records failed rather than
publishing a zero-coverage review. When summary posting is enabled, failed
sub-skills also appear in a distinct incomplete-coverage section rather than
being presented as skipped or as successful zero-finding reviews.
- A caller-provided
config_path, resolved from the target repository. - Individual workflow inputs such as
bcquality_repoandbcquality_ref, which override the selected config through environment variables. - When neither is supplied, the engine-owned
agents/ALReviewAgent/bcquality.config.yaml, whosebcquality.refis the reproducible source-of-truth pin.
A BCQuality dependency update changes both bcquality.ref to a reviewed commit
and bcquality.version to that content release. The latter changes Z in the
engine's X.Y.Z tag, causing the version workflow to publish the newly vouched
combination and move latest.
For additive producer/consumer contract changes, roll out producer-first:
- Merge and release the BCQuality producer change.
- Merge and release the backward-compatible engine consumer change.
- Update
agents/ALReviewAgent/bcquality.config.yamlto the BCQuality release commit that contains the producer change, updatebcquality.versionwith it, and verify the workflow's resolvedbcquality_shaplus representative rendered output.
Do not pin the engine to an unmerged BCQuality pull-request commit.
- The review job is read-only. It runs the tool-enabled Copilot CLI over untrusted PR-diff content and therefore never holds a write token.
- The publish job holds
issues/pull-requests: writebut never runs the model; it only posts findings saved as an artifact by the review job. - Both jobs check out with
persist-credentials: falseso a successful prompt-injection cannot exfiltrate a git token from.git/config. - BCQuality is cloned and filtered before the model runs. Point
bcquality.repoonly at a trusted source and pinbcquality.refto a reviewed commit — a compromised fork can embed prompt-injection payloads.
The orchestrator is entirely environment-variable driven and supports a
single-process mode (REVIEW_PHASE=all) that generates and posts in one pass —
used for local development and offline evaluation (e.g. BC-Bench). Provide a
BCQuality checkout via BCQUALITY_ROOT, the repo under review via
REVIEW_WORKSPACE, and point BCQUALITY_CONFIG_PATH at a policy config; then
invoke agents/ALReviewAgent/scripts/Invoke-CopilotPRReview.ps1.
In the reusable two-job workflow, BCQUALITY_ROOT is required only by the
generate phase. The checkout's resolved 40-character SHA is passed to the
checkout-free post phase as BCQUALITY_SHA; post requires that value and
publishes only the generated artifact.
The reusable workflow installs the exact copilot_cli_version input (default:
1.0.88), never latest. Direct callers must continue to set
COPILOT_REVIEW_CLI_VERSION explicitly. Before any model process starts, the
engine resolves the same executable used for leaf and root processes, runs
copilot --version, and extracts a strict semantic version from exactly one
official GitHub Copilot CLI <version>. banner. The requested pin must exactly
equal that startup probe and must have an explicit entry in
copilot-cli-compatibility.psd1.
| Exact startup version | OTel cli_version contract |
|---|---|
1.0.83 |
Required and must exactly equal the startup probe. |
1.0.88 |
May be absent; when present, it must exactly equal the startup probe. |
| Any other version | Rejected before model invocation as not compatibility-validated. |
Model identity, complete token usage, valid telemetry records, and each
process's requested-model contract remain strict for every supported release.
_run-manifest.json remains schema version 1:
configuration.copilot_cli_version is the startup-probed authoritative runtime
version. The caller/workflow pin is validated as exactly equal to that value
before any model process starts, so the existing v1 manifest shape needs no
second requested-version property.
To adopt a new CLI release, run the candidate exact version in a non-production compatibility canary; verify the executable probe and root/leaf model, usage, and telemetry contracts; then add an exact policy entry with its tested OTel behavior. Only after that change is reviewed and released should the reusable-workflow default be bumped. Production callers must remain pinned to a released engine SHA and an exact validated CLI version; no scheduled workflow moves either pin automatically.
Each generate/all run also writes _run-metrics.json to REVIEW_OUTPUT_DIR.
Schema version 1 has one 18-field shape and two metrics_source values:
copilot-cli-otel for executed reviews and not-applicable when the local
preflight skips a diff with no AL files. The latter reports exact zero use;
provider-dependent reasoning_tokens and legacy premium_requests remain null.
For executed reviews, prompt_tokens, completion_tokens, and nullable
reasoning_tokens sum the corresponding gen_ai.usage.* values on every raw
chat span, including nested agents, failures, and retries; total_tokens is
input plus output because reasoning is a subtype of output. api_calls counts
those spans rather than conversation turns, and wall_time_seconds covers the
full engine phase. cached_tokens and cache_creation_tokens are nullable when
the provider does not expose them. ai_credits is the exact sum of
github.copilot.nano_aiu divided by 1 billion, and premium_requests is the
exact sum of the legacy premium-request multiplier in github.copilot.cost;
either total is null unless every counted span exposes its source attribute.
usage_complete is false when any counted request lacks input/output usage.
Copilot CLI sub-agents launched through the task tool emit chat spans
without usage attributes, so every leaf and root review process runs with
--excluded-tools task; delegation would otherwise leave usage incomplete and
fail the review closed.
Invalid JSON lines and chat spans with invalid numeric/status attributes are
ignored independently and counted in malformed_records.
The OTel contract and attribute names were validated end to end with an isolated,
auto-update-disabled Copilot CLI 1.0.79 invocation. The raw JSONL is created
under the system temporary directory, message-content capture is disabled, and
the file is deleted immediately after harvesting; it is never copied to
REVIEW_OUTPUT_DIR. No console or transcript text is parsed.
This project is licensed under the MIT License.