Skip to content

Latest commit

 

History

History
381 lines (317 loc) · 19.8 KB

File metadata and controls

381 lines (317 loc) · 19.8 KB

Contributing to AgentFlow

Development setup

Use the Quick start in README.md and choose the setup script that matches your shell:

  • PowerShell: . .\scripts\setup.ps1
  • macOS / Linux: source ./scripts/setup.sh

Those scripts create the quick demo environment (.[dev] plus ./sdk). For workflow-faithful installs, use the canonical dependency profiles declared in pyproject.toml under [tool.agentflow.dependency-profiles]:

Profile Install contract Used by
runtime pip install -e . local serving/runtime-only paths
dev-tools pip install -e ".[dev]" lint, schema-check, host-side e2e, staging, backup
test pip install -e ".[dev,cloud]" test-integration, chaos, mutation
test-integrations pip install -e ".[dev,cloud]" + pip install -e "./sdk" + pip install -e "./integrations[mcp]" test-unit, local make setup
load pip install -e ".[load,cloud]" load-test
perf pip install -e ".[dev,load,cloud]" perf-check, perf-baseline, perf-smoke
contract pip install -e ".[dev,cloud,contract]" contract workflow

For the fastest local loop, use make demo. For a production-shaped local stack with observability, use make stack-prod-shaped-local (a demo, not a production recipe -- see docs/deployment.md).

Running tests

Release verification slice:

python -m pip install -e ".[dev,cloud]"
python -m pip install -e "./sdk"
python -m pip install -e "./integrations[mcp]"
python -m pytest tests/unit tests/integration tests/sdk -v

Additional suites when your change touches those areas:

python -m pip install -e ".[dev,cloud,contract]"
python -m pytest tests/contract tests/property tests/chaos tests/e2e -v
python -m pip install -e ".[dev,load,cloud]"
python scripts/run_benchmark.py
cd sdk-ts && npm test

The root integrations extra is intentionally not the repo test profile. Use ./integrations[mcp] when you need LangChain, LlamaIndex, and MCP coverage together.

After the package-identity split, pip show agentflow refers to the Python SDK and pip show agentflow-runtime refers to the root runtime repo metadata.

Review governance

Changes reach main through a pull request. Routine direct pushes to main are prohibited. A pull request must have at least one approval, and changes to the security, release, CI/CD, Helm, and infrastructure surfaces listed in .github/CODEOWNERS must also have Code Owner approval. The 15 required machine checks remain mandatory in addition to human review; they do not replace it.

This is the intended policy, not the current GitHub enforcement state. Adding CODEOWNERS does not activate branch protection by itself. The repository owner must update the protection rule or ruleset for main to:

  1. require a pull request before merging with at least one approving review;
  2. require review from Code Owners;
  3. keep all 15 existing required status checks enabled;
  4. prevent direct pushes and apply the rule to administrators; or restrict the only bypass to a named break-glass role whose use is recorded in the organization audit log, linked to an incident or change ticket, and reviewed after the event.

Until those settings are enabled, review governance remains open and must not be reported as enforced.

Before submitting a PR

  1. Tests pass:
make test
  1. Security diff is clean:
mkdir -p .artifacts/security
bandit -r src sdk --ini .bandit --severity-level medium -f json -o .artifacts/security/bandit-current.json
python scripts/bandit_diff.py .bandit-baseline.json .artifacts/security/bandit-current.json
  1. Benchmark does not regress past the release gate:
python scripts/run_benchmark.py
python scripts/check_performance.py --baseline docs/benchmark-baseline.json --current .artifacts/benchmark/current.json --max-regress 20
  1. Contracts are still in sync:
python scripts/generate_contracts.py --check
python scripts/export_openapi.py --check
python scripts/export_sdk_capabilities.py --check
python scripts/export_quality_reference.py --check

After changing an API route or schema, run python scripts/export_openapi.py and commit all three outputs (docs/openapi.json plus both files under docs/agent-tools/). Do not edit one generated output independently.

After changing [sdk] claims or either public SDK client surface, run python scripts/export_sdk_capabilities.py and commit docs/sdk-capabilities.md. The --check form reproduces the tracked output; the project-claims validator also checks that every declared method exists.

After changing [quality] claims, run python scripts/export_quality_reference.py and commit docs/quality.md. Host-specific reports belong to the ignored default output of python scripts/quality_report.py; do not use that collector to overwrite the tracked current reference.

python scripts/dora_metrics.py --days 30 writes the host-, time-, and GitHub-history-dependent DORA JSON to ignored .artifacts/dora/dora-report.json. Relative outputs resolve from the project root. Do not leave dora-report.json, dora-summary.md, or dora-comment.md in the repository root; the weekly/PR workflow also keeps those working files under .artifacts/dora/. This is runtime evidence, not production acceptance or a byte-regenerated reference. Promote a reviewed snapshot only under a new date-stamped identity with source SHA, window/branch, data sources, exact command/configuration, host/runtime, and artifact hash provenance.

python scripts/chaos_report.py reads ignored .artifacts/chaos/chaos-report.json by default and may write .artifacts/chaos/chaos-summary.json plus .artifacts/chaos/chaos-summary.md. Relative --input, --output, and --markdown paths resolve from the project root; optional output parents are created. Do not leave chaos-report.json, chaos-summary.json, or chaos-summary.md in the repository root; the chaos workflow already keeps those working files under .artifacts/chaos/. This is host/time/test-run-dependent runtime evidence, not production acceptance or a byte-regenerated reference. Promote a reviewed snapshot only under a new date-stamped identity with source SHA, scenario/configuration, host/runtime, exact command, result counts, and artifact hashes.

python scripts/mutation_report.py writes ignored mutation JSON and work files under .artifacts/mutation/ (mutmut-cicd-stats.json and per-module .meta files). Relative --results-dir resolves from the project root, not the caller CWD, and every destination under docs/ is rejected before mutmut runs. The weekly mutation workflow uploads .artifacts/mutation/. These files are replaceable runtime artifacts, not reviewed evidence or production acceptance. Promote a reviewed snapshot only under a new date-stamped identity with provenance.

python scripts/mutation_local.py --module <target> measures one module of that same gate on this machine, so a mutation score is available before you push instead of only after the weekly workflow. --list-modules prints the targets; the score, the mutant population and the surviving mutant names match the CI run for the same commit — up to pytest's internal-error exits, which scripts/mutation_report.py counts as kills (exit 3) while this driver refuses to score them at all — because the driver reads its targets from scripts/mutation_report.MODULE_TARGETS, builds its workspace with prepare_workspace, and generates mutants with mutmut's own engine (mutmut.mutation.file_mutation). It needs mutmut installed (it is in the dev extra) but never invokes the mutmut CLI: mutants are executed as plain pytest subprocesses selected through MUTANT_UNDER_TEST, which is what makes this work on native Windows, where mutmut run exits at import time. Expect minutes, not seconds — one pytest process per mutant, --jobs in parallel. The driver exits 1 when the module is below its threshold or when any mutant got no verdict (only pytest exit 0 = survived and 1 = killed are verdicts; anything else is a harness failure, never a kill). A mutant that comes back without a verdict — usually a timeout from running --jobs of them at once — is retried once serially with a longer timeout before it is reported that way, so the score does not move with the machine's load. Its workspace and JSON report live under the OS temp directory, outside the repository; the workspace is reused across runs only when it is stamped with the same root, module, module source, materialized top-level package tree, target tests and pyproject.toml (which prepare_workspace always renders into the workspace as a real file, carrying pytest addopts, filterwarnings and [tool.mutmut]), and is rebuilt otherwise, so a second run never reports the first one's copy of those sources. The stamp does not cover the trees prepare_workspace normally symlinks — src/, sdk/, config/, scripts/ and the rest of tests/ beyond the target's own test files — which it copies instead where the OS refuses symlinks; on such a machine, pass a fresh --workspace after editing them. A --workspace is emptied on rebuild, so one that is a checkout — this repository, anything inside the tree the sources come from, or any directory holding a .git — is refused instead, and only a directory carrying the driver's own marker is ever cleared. --root points it at another checkout; --only re-runs named mutants, written either bare or exactly as the report prints them (<module>.<mutant>).

python scripts/evaluate_trivy_policy.py writes ignored Trivy policy summaries under .artifacts/trivy/. Relative --report, --waivers, and --output paths resolve from the project root, not the caller CWD, and every destination under docs/ is rejected before the report is read. The security workflow keeps SBOM, JSON, SARIF, policy-summary, and IaC working files under .artifacts/trivy/. These files are replaceable runtime/CI artifacts, not reviewed evidence or production acceptance. Promote a reviewed snapshot only under a new date-stamped identity with provenance.

The Scorecard workflow writes one ignored per-run SARIF working copy to .artifacts/scorecard/results.sarif. That local/uploaded SARIF is not reviewed evidence, a penetration-test attestation, or production acceptance. The Code scanning upload and public OpenSSF registry result remain the channel outputs. Promote a reviewed snapshot only under a new date-stamped identity with source SHA, workflow run, tool/action version, exact configuration, and hash provenance.

The security workflow keeps its dependency-scan working files under ignored .artifacts/security/. The Bandit job writes bandit-current.json there and diffs it against the tracked .bandit-baseline.json, which is the only reviewed input; the local Bandit command above uses the same path. The Safety job resolves its requirement buckets, resolver virtualenvs, and the vulnerable-pin regression probe under .artifacts/security/safety/. The pip-audit job exports the full locked profile set to .artifacts/security/pip-audit/requirements-all-profiles.txt; its production step still reads the tracked requirements-docker.lock directly. Do not leave scanner output in the repository root or .tmp/. These files are replaceable per-run working copies, not reviewed evidence, a dependency-compatibility attestation, or production acceptance. Promote a reviewed scan only under a new date-stamped identity with source SHA, workflow run, scanner versions, exact command and configuration, outcome, and hash provenance.

The Terraform workflow keeps its plan and apply jobs disabled (if: false) until AWS is explicitly reintroduced, but their artifact contract is fixed now: the plan job writes the binary plan to ignored .artifacts/terraform/tfplan (addressed through $GITHUB_WORKSPACE because both steps run from infrastructure/terraform/), uploads it as terraform-plan-<environment>, and the apply job downloads the same path. Never write a plan file next to the configuration or commit one: plan files embed resolved variable values. The plan file is a replaceable per-run working copy, not reviewed evidence, OIDC/apply evidence, or production acceptance. Promote a reviewed plan only under a new date-stamped identity with source SHA, workflow run, Terraform/action versions, tfvars identity, exact command, outcome, and hash provenance.

python scripts/profile_entity.py --entity-type <type> --entity-id <id> writes the quick entity-latency runtime result to ignored .artifacts/perf-smoke/entity-profile.json. Relative outputs resolve from the project root, and the harness refuses to write anywhere under docs/perf/. Promote only a reviewed run under a new date-stamped identity with its host/runtime, source SHA, exact command, sample counts, and profile write-up.

python scripts/run_benchmark.py writes its host- and time-dependent report to .artifacts/benchmark/benchmark.md and its JSON metrics to .artifacts/benchmark/current.json. These runtime outputs have no byte-drift check and must not replace docs/perf/load-benchmark-latest.md or an archived snapshot. They also must not replace docs/benchmark-baseline.json, which is a reviewed gate-policy input rather than generated runtime output. CI compares the fresh JSON metrics with that tracked gate baseline; promote evidence only under a date-stamped name with its run provenance.

The GitHub ARM workflow writes the same host-dependent pair plus host metadata to ignored .artifacts/benchmark/arm-*. Those runtime artifacts must not replace the four immutable 2026-06-05 files under docs/perf/. The harness rejects those reviewed tracked paths as runtime output. Promote a reviewed ARM run only under a new date-stamped identity with source, host/runtime, exact command/configuration, sample/threshold information, and artifact hashes.

python tests/load/run_load_test.py writes the Locust p99 CI-smoke CSV prefix to .artifacts/load/results and JSON metrics to .artifacts/load/results.json. Relative outputs resolve from the project root, and the runner refuses destinations under docs/perf/ or tests/load/ before seed or Locust work. make load-test invokes this runner with the localhost default and the 50 users / 10 spawn-rate / 60-second profile. Compare a run with python scripts/check_performance.py --baseline docs/benchmark-baseline.json --current .artifacts/load/results.json. This is host- and time-dependent CI-smoke runtime evidence, not a byte-regenerated tracked reference, production SLA, full-load benchmark, or acceptance. Promote a reviewed result only under a new date-stamped identity with provenance.

To compare repeated local runs, append one results file with python scripts/record_perf_history.py --results .artifacts/benchmark/current.json, then run python scripts/plot_perf_history.py. The commands own .artifacts/perf-history/history.json, history.html, and optional history.png; they refuse to overwrite the retired tracked history or write plots under docs/. CI does not persist this history across runners, so do not cite the local trend as continuous CI or release evidence.

python scripts/benchmark_freshness.py writes the in-process demo report to .artifacts/freshness/freshness-benchmark.md and machine-readable results to .artifacts/freshness/current.json. Do not overwrite the tracked docs/perf/freshness-benchmark.md lifecycle page or its archived snapshot. Promote a meaningful run only under a date-stamped name with its JSON companion and exact run, host, and source provenance.

python scripts/benchmark_freshness_realpath.py writes the Kafka → Flink streaming-hop result to .artifacts/freshness/realpath-current.json. Run its Kafka/Flink prerequisites on deproject-mac, not the Windows host. The driver refuses to overwrite the immutable docs/perf/freshness-realpath-2026-06-30.md record; promote a reviewed run only under a new date-stamped evidence identity with exact source, host, runtime, command, configuration, sample count, miss count, and JSON hash.

python scripts/benchmark_freshness_e2e.py writes the S8 real-path report to .artifacts/freshness/e2e-realpath.md and machine-readable results to .artifacts/freshness/e2e-realpath-current.json. Run its Kafka/Flink/bridge/ ClickHouse/Redis/API prerequisites on deproject-mac, not the Windows host. Do not overwrite the tracked docs/perf/freshness-e2e-realpath.md lifecycle page or its archived 2026-07-09 snapshot; promote only date-stamped evidence with exact run, host, source, configuration, and JSON-companion provenance.

python scripts/benchmark_throughput_realpath.py writes the real-path report to .artifacts/throughput/realpath-current.md and machine-readable results to .artifacts/throughput/realpath-current.json. Run its Kafka/Flink/bridge/ ClickHouse prerequisites on deproject-mac, not the Windows development host. Do not overwrite the tracked docs/perf/throughput-realpath.md lifecycle page or its archived S10 baseline; promote only date-stamped evidence with exact run, host, source, and configuration provenance.

python scripts/benchmark_scale_own_data.py writes the S13 own-data scale reports to .artifacts/scale/own-data-current.md and .artifacts/scale/own-data-current.json. Run its live ClickHouse workload on deproject-mac, not the Windows host. The driver refuses to overwrite docs/perf/scale-own-data-2026-07-11.md; promote a reviewed run only under a new date-stamped identity with exact source, host/runtime, command, configuration, volume/check results, and Markdown/JSON hashes.

python scripts/perf/auth_bench.py writes its host-dependent legacy-path microbenchmark to .artifacts/perf/auth-bench-current.md. Run the full bcrypt workload on deproject-mac, not the Windows development host. The driver uses explicit legacy bcrypt semantics and refuses to overwrite docs/perf/auth-bench.md or the immutable 2026-05-26 record. Promote a reviewed run only under a new date-stamped identity with exact source, host/power, Python/dependency, command/configuration, sample-count, boundary, and report-hash provenance.

python -m scripts.run_nl_sql_eval writes the direct-translator result to .artifacts/nl-sql-eval/current.md. Relative outputs resolve from the project root, and the command rejects every path under docs/perf/ before running the evaluation. In particular, it cannot overwrite docs/perf/nl-sql-eval-2026-07-01.md or docs/perf/nl-sql-eval-sonnet5-2026-07-01.md. The rule-based path uses the fixed in-memory DuckDB demo set; the opt-in LLM path is live and non-deterministic. Neither is the served /query path, a production benchmark, an SLA, or acceptance. Promote a reviewed result only under a new date-stamped identity with source, host/runtime, engine/model, exact command/configuration, and report-hash provenance.

Dependabot pip PRs and uv.lock

Dependabot bumps grouped pip dependencies in pyproject.toml without regenerating uv.lock, so the lock-check gate (uv lock --check) fails on every such PR. Bring the lockfile along by hand:

gh pr checkout <pr-number>
uv lock
git add uv.lock && git commit -m "chore(deps): regenerate uv.lock for the group bump"
git push
gh pr update-branch <pr-number>   # branch protection requires up-to-date branches

Then let CI finish and merge as usual. GitHub Actions and npm bumps do not need this — only the pip ecosystem is locked with uv.

Architecture decisions

Significant design changes should include an ADR in docs/decisions/.

Start with:

Documentation expectations

If you change the HTTP surface or operational behavior, update the matching docs:

  • docs/api-reference.md
  • docs/architecture.md
  • docs/runbook.md
  • docs/security-audit.md when the control surface changes

Commit conventions

Use conventional commit prefixes:

  • feat:
  • fix:
  • docs:
  • chore:
  • refactor:
  • test: