Use the Quick start in README.md and choose the setup script that matches your shell:
- PowerShell:
. .\scripts\setup.ps1 - macOS / Linux:
source ./scripts/setup.sh
Those scripts create the quick demo environment (.[dev] plus ./sdk). For workflow-faithful installs, use the canonical dependency profiles declared in pyproject.toml under [tool.agentflow.dependency-profiles]:
| Profile | Install contract | Used by |
|---|---|---|
runtime |
pip install -e . |
local serving/runtime-only paths |
dev-tools |
pip install -e ".[dev]" |
lint, schema-check, host-side e2e, staging, backup |
test |
pip install -e ".[dev,cloud]" |
test-integration, chaos, mutation |
test-integrations |
pip install -e ".[dev,cloud]" + pip install -e "./sdk" + pip install -e "./integrations[mcp]" |
test-unit, local make setup |
load |
pip install -e ".[load,cloud]" |
load-test |
perf |
pip install -e ".[dev,load,cloud]" |
perf-check, perf-baseline, perf-smoke |
contract |
pip install -e ".[dev,cloud,contract]" |
contract workflow |
For the fastest local loop, use make demo. For a production-shaped local stack with observability, use make stack-prod-shaped-local (a demo, not a production recipe -- see docs/deployment.md).
Release verification slice:
python -m pip install -e ".[dev,cloud]"
python -m pip install -e "./sdk"
python -m pip install -e "./integrations[mcp]"
python -m pytest tests/unit tests/integration tests/sdk -vAdditional suites when your change touches those areas:
python -m pip install -e ".[dev,cloud,contract]"
python -m pytest tests/contract tests/property tests/chaos tests/e2e -v
python -m pip install -e ".[dev,load,cloud]"
python scripts/run_benchmark.py
cd sdk-ts && npm testThe root integrations extra is intentionally not the repo test profile. Use ./integrations[mcp] when you need LangChain, LlamaIndex, and MCP coverage together.
After the package-identity split, pip show agentflow refers to the Python SDK and pip show agentflow-runtime refers to the root runtime repo metadata.
Changes reach main through a pull request. Routine direct pushes to main
are prohibited. A pull request must have at least one approval, and changes to
the security, release, CI/CD, Helm, and infrastructure surfaces listed in
.github/CODEOWNERS must also have Code Owner approval. The 15 required
machine checks remain mandatory in addition to human review; they do not
replace it.
This is the intended policy, not the current GitHub enforcement state. Adding
CODEOWNERS does not activate branch protection by itself. The repository
owner must update the protection rule or ruleset for main to:
- require a pull request before merging with at least one approving review;
- require review from Code Owners;
- keep all 15 existing required status checks enabled;
- prevent direct pushes and apply the rule to administrators; or restrict the only bypass to a named break-glass role whose use is recorded in the organization audit log, linked to an incident or change ticket, and reviewed after the event.
Until those settings are enabled, review governance remains open and must not be reported as enforced.
- Tests pass:
make test- Security diff is clean:
mkdir -p .artifacts/security
bandit -r src sdk --ini .bandit --severity-level medium -f json -o .artifacts/security/bandit-current.json
python scripts/bandit_diff.py .bandit-baseline.json .artifacts/security/bandit-current.json- Benchmark does not regress past the release gate:
python scripts/run_benchmark.py
python scripts/check_performance.py --baseline docs/benchmark-baseline.json --current .artifacts/benchmark/current.json --max-regress 20- Contracts are still in sync:
python scripts/generate_contracts.py --check
python scripts/export_openapi.py --check
python scripts/export_sdk_capabilities.py --check
python scripts/export_quality_reference.py --checkAfter changing an API route or schema, run python scripts/export_openapi.py
and commit all three outputs (docs/openapi.json plus both files under
docs/agent-tools/). Do not edit one generated output independently.
After changing [sdk] claims or either public SDK client surface, run
python scripts/export_sdk_capabilities.py and commit
docs/sdk-capabilities.md. The --check form reproduces the tracked output;
the project-claims validator also checks that every declared method exists.
After changing [quality] claims, run
python scripts/export_quality_reference.py and commit docs/quality.md.
Host-specific reports belong to the ignored default output of
python scripts/quality_report.py; do not use that collector to overwrite the
tracked current reference.
python scripts/dora_metrics.py --days 30 writes the host-, time-, and
GitHub-history-dependent DORA JSON to ignored
.artifacts/dora/dora-report.json. Relative outputs resolve from the project
root. Do not leave dora-report.json, dora-summary.md, or dora-comment.md
in the repository root; the weekly/PR workflow also keeps those working files
under .artifacts/dora/. This is runtime evidence, not production acceptance
or a byte-regenerated reference. Promote a reviewed snapshot only under a new
date-stamped identity with source SHA, window/branch, data sources, exact
command/configuration, host/runtime, and artifact hash provenance.
python scripts/chaos_report.py reads ignored
.artifacts/chaos/chaos-report.json by default and may write
.artifacts/chaos/chaos-summary.json plus .artifacts/chaos/chaos-summary.md.
Relative --input, --output, and --markdown paths resolve from the project
root; optional output parents are created. Do not leave chaos-report.json,
chaos-summary.json, or chaos-summary.md in the repository root; the chaos
workflow already keeps those working files under .artifacts/chaos/. This is
host/time/test-run-dependent runtime evidence, not production acceptance or a
byte-regenerated reference. Promote a reviewed snapshot only under a new
date-stamped identity with source SHA, scenario/configuration, host/runtime,
exact command, result counts, and artifact hashes.
python scripts/mutation_report.py writes ignored mutation JSON and work files
under .artifacts/mutation/ (mutmut-cicd-stats.json and per-module .meta
files). Relative --results-dir resolves from the project root, not the caller
CWD, and every destination under docs/ is rejected before mutmut runs. The
weekly mutation workflow uploads .artifacts/mutation/. These files are
replaceable runtime artifacts, not reviewed evidence or production acceptance.
Promote a reviewed snapshot only under a new date-stamped identity with
provenance.
python scripts/mutation_local.py --module <target> measures one module of
that same gate on this machine, so a mutation score is available before you
push instead of only after the weekly workflow. --list-modules prints the
targets; the score, the mutant population and the surviving mutant names match
the CI run for the same commit — up to pytest's internal-error exits, which
scripts/mutation_report.py counts as kills (exit 3) while this driver refuses
to score them at all — because the driver reads its targets from
scripts/mutation_report.MODULE_TARGETS, builds its workspace with
prepare_workspace, and generates mutants with mutmut's own engine
(mutmut.mutation.file_mutation). It needs mutmut installed (it is in the
dev extra) but never invokes the mutmut CLI: mutants are executed as plain
pytest subprocesses selected through MUTANT_UNDER_TEST, which is what makes
this work on native Windows, where mutmut run exits at import time. Expect
minutes, not seconds — one pytest process per mutant, --jobs in parallel.
The driver exits 1 when the module is below its threshold or when any mutant
got no verdict (only pytest exit 0 = survived and 1 = killed are verdicts;
anything else is a harness failure, never a kill). A mutant that comes back
without a verdict — usually a timeout from running --jobs of them at once —
is retried once serially with a longer timeout before it is reported that way,
so the score does not move with the machine's load. Its workspace and JSON
report live under the OS temp directory, outside the repository; the workspace
is reused across runs only when it is stamped with the same root, module,
module source, materialized top-level package tree, target tests and
pyproject.toml (which prepare_workspace always renders into the workspace as
a real file, carrying pytest addopts, filterwarnings and [tool.mutmut]), and
is rebuilt otherwise, so a second run never reports the first one's copy of
those sources. The stamp does not cover the trees prepare_workspace normally
symlinks — src/, sdk/, config/, scripts/ and the rest of tests/
beyond the target's own test files — which it copies instead where the OS
refuses symlinks; on such a machine, pass a fresh --workspace after editing
them. A --workspace is emptied on rebuild,
so one that is a checkout — this repository, anything inside the tree the
sources come from, or any directory holding a .git — is refused instead, and
only a directory carrying the driver's own marker is ever cleared. --root
points it at another checkout; --only re-runs named mutants, written either
bare or exactly as the report prints them (<module>.<mutant>).
python scripts/evaluate_trivy_policy.py writes ignored Trivy policy
summaries under .artifacts/trivy/. Relative --report, --waivers, and
--output paths resolve from the project root, not the caller CWD, and every
destination under docs/ is rejected before the report is read. The security
workflow keeps SBOM, JSON, SARIF, policy-summary, and IaC working files under
.artifacts/trivy/. These files are replaceable runtime/CI artifacts, not
reviewed evidence or production acceptance. Promote a reviewed snapshot only
under a new date-stamped identity with provenance.
The Scorecard workflow writes one ignored per-run SARIF working copy to
.artifacts/scorecard/results.sarif. That local/uploaded SARIF is not
reviewed evidence, a penetration-test attestation, or production acceptance.
The Code scanning upload and public OpenSSF registry result remain the
channel outputs. Promote a reviewed snapshot only under a new date-stamped
identity with source SHA, workflow run, tool/action version, exact
configuration, and hash provenance.
The security workflow keeps its dependency-scan working files under ignored
.artifacts/security/. The Bandit job writes bandit-current.json there and
diffs it against the tracked .bandit-baseline.json, which is the only
reviewed input; the local Bandit command above uses the same path. The Safety
job resolves its requirement buckets, resolver virtualenvs, and the
vulnerable-pin regression probe under .artifacts/security/safety/. The
pip-audit job exports the full locked profile set to
.artifacts/security/pip-audit/requirements-all-profiles.txt; its production
step still reads the tracked requirements-docker.lock directly. Do not leave
scanner output in the repository root or .tmp/. These files are replaceable
per-run working copies, not reviewed evidence, a dependency-compatibility
attestation, or production acceptance. Promote a reviewed scan only under a
new date-stamped identity with source SHA, workflow run, scanner versions,
exact command and configuration, outcome, and hash provenance.
The Terraform workflow keeps its plan and apply jobs disabled (if: false)
until AWS is explicitly reintroduced, but their artifact contract is fixed now:
the plan job writes the binary plan to ignored .artifacts/terraform/tfplan
(addressed through $GITHUB_WORKSPACE because both steps run from
infrastructure/terraform/), uploads it as terraform-plan-<environment>, and
the apply job downloads the same path. Never write a plan file next to the
configuration or commit one: plan files embed resolved variable values. The
plan file is a replaceable per-run working copy, not reviewed evidence,
OIDC/apply evidence, or production acceptance. Promote a reviewed plan only
under a new date-stamped identity with source SHA, workflow run,
Terraform/action versions, tfvars identity, exact command, outcome, and hash
provenance.
python scripts/profile_entity.py --entity-type <type> --entity-id <id> writes
the quick entity-latency runtime result to ignored
.artifacts/perf-smoke/entity-profile.json. Relative outputs resolve from the
project root, and the harness refuses to write anywhere under docs/perf/.
Promote only a reviewed run under a new date-stamped identity with its
host/runtime, source SHA, exact command, sample counts, and profile write-up.
python scripts/run_benchmark.py writes its host- and time-dependent report to
.artifacts/benchmark/benchmark.md and its JSON metrics to
.artifacts/benchmark/current.json. These runtime outputs have no byte-drift
check and must not replace docs/perf/load-benchmark-latest.md or an archived
snapshot. They also must not replace docs/benchmark-baseline.json, which is a
reviewed gate-policy input rather than generated runtime output. CI compares
the fresh JSON metrics with that tracked gate baseline; promote evidence only
under a date-stamped name with its run provenance.
The GitHub ARM workflow writes the same host-dependent pair plus host metadata
to ignored .artifacts/benchmark/arm-*. Those runtime artifacts must not
replace the four immutable 2026-06-05 files under docs/perf/. The harness
rejects those reviewed tracked paths as runtime output. Promote a reviewed ARM
run only under a new date-stamped identity with source, host/runtime, exact
command/configuration, sample/threshold information, and artifact hashes.
python tests/load/run_load_test.py writes the Locust p99 CI-smoke CSV prefix
to .artifacts/load/results and JSON metrics to .artifacts/load/results.json.
Relative outputs resolve from the project root, and the runner refuses
destinations under docs/perf/ or tests/load/ before seed or Locust work.
make load-test invokes this runner with the localhost default and the
50 users / 10 spawn-rate / 60-second profile. Compare a run with
python scripts/check_performance.py --baseline docs/benchmark-baseline.json --current .artifacts/load/results.json. This is host- and time-dependent
CI-smoke runtime evidence, not a byte-regenerated tracked reference, production
SLA, full-load benchmark, or acceptance. Promote a reviewed result only under
a new date-stamped identity with provenance.
To compare repeated local runs, append one results file with
python scripts/record_perf_history.py --results .artifacts/benchmark/current.json,
then run python scripts/plot_perf_history.py. The commands own
.artifacts/perf-history/history.json, history.html, and optional
history.png; they refuse to overwrite the retired tracked history or write
plots under docs/. CI does not persist this history across runners, so do not
cite the local trend as continuous CI or release evidence.
python scripts/benchmark_freshness.py writes the in-process demo report to
.artifacts/freshness/freshness-benchmark.md and machine-readable results to
.artifacts/freshness/current.json. Do not overwrite the tracked
docs/perf/freshness-benchmark.md lifecycle page or its archived snapshot.
Promote a meaningful run only under a date-stamped name with its JSON companion
and exact run, host, and source provenance.
python scripts/benchmark_freshness_realpath.py writes the Kafka → Flink
streaming-hop result to .artifacts/freshness/realpath-current.json. Run its
Kafka/Flink prerequisites on deproject-mac, not the Windows host. The driver
refuses to overwrite the immutable
docs/perf/freshness-realpath-2026-06-30.md record; promote a reviewed run only
under a new date-stamped evidence identity with exact source, host, runtime,
command, configuration, sample count, miss count, and JSON hash.
python scripts/benchmark_freshness_e2e.py writes the S8 real-path report to
.artifacts/freshness/e2e-realpath.md and machine-readable results to
.artifacts/freshness/e2e-realpath-current.json. Run its Kafka/Flink/bridge/
ClickHouse/Redis/API prerequisites on deproject-mac, not the Windows host.
Do not overwrite the tracked docs/perf/freshness-e2e-realpath.md lifecycle
page or its archived 2026-07-09 snapshot; promote only date-stamped evidence
with exact run, host, source, configuration, and JSON-companion provenance.
python scripts/benchmark_throughput_realpath.py writes the real-path report
to .artifacts/throughput/realpath-current.md and machine-readable results to
.artifacts/throughput/realpath-current.json. Run its Kafka/Flink/bridge/
ClickHouse prerequisites on deproject-mac, not the Windows development host.
Do not overwrite the tracked docs/perf/throughput-realpath.md lifecycle page
or its archived S10 baseline; promote only date-stamped evidence with exact
run, host, source, and configuration provenance.
python scripts/benchmark_scale_own_data.py writes the S13 own-data scale
reports to .artifacts/scale/own-data-current.md and
.artifacts/scale/own-data-current.json. Run its live ClickHouse workload on
deproject-mac, not the Windows host. The driver refuses to overwrite
docs/perf/scale-own-data-2026-07-11.md; promote a reviewed run only under a
new date-stamped identity with exact source, host/runtime, command,
configuration, volume/check results, and Markdown/JSON hashes.
python scripts/perf/auth_bench.py writes its host-dependent legacy-path
microbenchmark to .artifacts/perf/auth-bench-current.md. Run the full bcrypt
workload on deproject-mac, not the Windows development host. The driver uses
explicit legacy bcrypt semantics and refuses to overwrite
docs/perf/auth-bench.md or the immutable 2026-05-26 record. Promote a reviewed
run only under a new date-stamped identity with exact source, host/power,
Python/dependency, command/configuration, sample-count, boundary, and report-hash
provenance.
python -m scripts.run_nl_sql_eval writes the direct-translator result to
.artifacts/nl-sql-eval/current.md. Relative outputs resolve from the project
root, and the command rejects every path under docs/perf/ before running the
evaluation. In particular, it cannot overwrite
docs/perf/nl-sql-eval-2026-07-01.md or
docs/perf/nl-sql-eval-sonnet5-2026-07-01.md. The rule-based path uses the
fixed in-memory DuckDB demo set; the opt-in LLM path is live and
non-deterministic. Neither is the served /query path, a production benchmark,
an SLA, or acceptance. Promote a reviewed result only under a new date-stamped
identity with source, host/runtime, engine/model, exact command/configuration,
and report-hash provenance.
Dependabot bumps grouped pip dependencies in pyproject.toml without
regenerating uv.lock, so the lock-check gate (uv lock --check) fails on
every such PR. Bring the lockfile along by hand:
gh pr checkout <pr-number>
uv lock
git add uv.lock && git commit -m "chore(deps): regenerate uv.lock for the group bump"
git push
gh pr update-branch <pr-number> # branch protection requires up-to-date branchesThen let CI finish and merge as usual. GitHub Actions and npm bumps do not need this — only the pip ecosystem is locked with uv.
Significant design changes should include an ADR in docs/decisions/.
Start with:
- docs/architecture.md
- docs/release-readiness.md
- existing ADRs under
docs/decisions/
If you change the HTTP surface or operational behavior, update the matching docs:
docs/api-reference.mddocs/architecture.mddocs/runbook.mddocs/security-audit.mdwhen the control surface changes
Use conventional commit prefixes:
feat:fix:docs:chore:refactor:test: