Repository navigation
Revival: test agent skills with BehaviorCI (from_skill_tests adapter) - #2
Conversation
Map an agent skill's tests.json acceptance cases onto @behavior tests: each case's input is handed to a user-supplied runner, expected.contains becomes the must_contain guardrail, and the output is snapshotted so a SKILL.md edit that keeps the required substrings but changes what the skill says still fails CI as semantic drift. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A german-b2b-outreach-email skill folder (SKILL.md + tests.json) plus a deterministic canned runner show the full flow; the examples double as dogfooding in CI's record/check step. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
README gains a 'Why BehaviorCI now' section (agent skills widen the prompt-regression surface) and a worked from_skill_tests() example; the docs site gains an 'Agent skills' page covering the tests.json mapping, failure output, options, and runner hygiene. Changelog updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Vouch AnalysisAnalysis 🛡️ Maintainer Gate — 2/4 evidence checks passedNote Some evidence checks did not pass (advisory for this author).
Warning Security or hallucination findings were detected in this pull request. 🚨 Security & Hallucinations
Note Dependency quality suggestions are advisory and do not block merge on their own. 🧹 Dependency Quality
🤖 AI System Disclosure: This analysis was generated by Vouch using deterministic AST parsing and registry verification (no LLM inference on this run). Learn more about how Vouch works. |
There was a problem hiding this comment.
Pull request overview
Adds first-class “agent skill” support to BehaviorCI by introducing an adapter that converts a skill folder’s tests.json cases into standard @behavior snapshot tests, along with documentation and runnable examples to dogfood the record/check workflow in CI.
Changes:
- Introduces
from_skill_tests()adapter to generate one@behaviortest pertests.jsoncase (withexpected.contains → must_containmapping). - Adds direct + pytester end-to-end tests plus a runnable example skill fixture under
tests/examples/. - Adds new docs page and updates README, MkDocs nav, and changelog to document skill testing.
Reviewed changes
Copilot reviewed 11 out of 11 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| tests/test_skills.py | Adds unit + pytester end-to-end coverage for from_skill_tests() generation/validation and record→check behavior. |
| tests/examples/test_skill.py | Provides a deterministic, runnable example showing how to wire from_skill_tests() into a test module. |
| tests/examples/skills/german-b2b-outreach-email/tests.json | Adds example acceptance cases (input + expected.contains) for the sample skill. |
| tests/examples/skills/german-b2b-outreach-email/SKILL.md | Adds example skill instructions used by the adapter example. |
| src/behaviorci/skills.py | Implements the new skill adapter (from_skill_tests) and validation of tests.json. |
| src/behaviorci/init.py | Exports from_skill_tests from the top-level package API. |
| README.md | Documents “Testing agent skills” and positions BehaviorCI’s value for skills/drift. |
| mkdocs.yml | Adds the new “Agent skills” page to the docs navigation. |
| docs/skills.md | Adds dedicated documentation for skill folder format, adapter usage, and failure output. |
| docs/index.md | Links the new Agent skills docs page from the docs landing page. |
| CHANGELOG.md | Records the new adapter/docs/example under Unreleased additions. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| skill_path = Path(skill_dir) | ||
| cases = _load_cases(skill_path / TESTS_FILENAME) | ||
|
|
||
| prefix = id_prefix or skill_path.resolve().name |
numpy>=2.5 ships type aliases in its .pyi (valid only for target >=3.12); under our 3.10 mypy target these are rejected as syntax errors, failing the lint job on a fresh dependency resolve even though our own code is unchanged. numpy is third-party (already under --ignore-missing-imports), so skip following into its stubs. Verified green on mypy 1.20.0 and 2.1.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Addresses review feedback: skill_path.resolve().name follows symlinks, so a symlinked skill folder produced test ids from the link target's name rather than the folder the caller passed. Use skill_path.name (falling back to resolve() only for '.'/'/'), so ids are stable and predictable across symlinked skill dirs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds from_skill_tests() to turn a skill folder's tests.json into @behavior snapshot tests, with docs + runnable example. Includes CI fixes: numpy PEP695-stub mypy override, and symlink-stable test ids. All checks green.
Prepared by an automated session for review before merge — not auto-merged, not published to PyPI.
What this adds (purely additive: +606 / -0)
from_skill_tests(skill_dir, run, threshold=0.85)insrc/behaviorci/skills.py— reads an agent skill folder'stests.json(expected.contains→must_contain) and emits one@behaviorsnapshot test per case, so aSKILL.mdedit that keeps the required substrings but changes what the skill says still fails CI as semantic drift.tests/examples/(SKILL.md + tests.json + a deterministic runner), picked up by the existing dogfood record/check CI step.docs/skills.md), README "Why BehaviorCI now" + skill-testing sections, CHANGELOG entry.Verification
mkdocs build --strictclean.Not done (left to the owner)
main, PyPI publish, and any launch post.🤖 Generated with Claude Code