Skip to content

Revival: test agent skills with BehaviorCI (from_skill_tests adapter) - #2

Merged
0-uddeshya-0 merged 5 commits into
mainfrom
revival
Jul 7, 2026
Merged

0-uddeshya-0 merged 5 commits into
mainfrom
revival

Conversation

@0-uddeshya-0

Copy link
Copy Markdown
Owner

Prepared by an automated session for review before merge — not auto-merged, not published to PyPI.

What this adds (purely additive: +606 / -0)

  • from_skill_tests(skill_dir, run, threshold=0.85) in src/behaviorci/skills.py — reads an agent skill folder's tests.json (expected.contains → must_contain) and emits one @behavior snapshot test per case, so a SKILL.md edit that keeps the required substrings but changes what the skill says still fails CI as semantic drift.
  • Worked runnable example under tests/examples/ (SKILL.md + tests.json + a deterministic runner), picked up by the existing dogfood record/check CI step.
  • Docs: new "Agent skills" page (docs/skills.md), README "Why BehaviorCI now" + skill-testing sections, CHANGELOG entry.
  • 15 new tests (generation, config mapping, validation errors, 3 pytester end-to-end).

Verification

  • Full suite: 115 passed (100 pre-existing + 15 new), independently re-run.
  • black / isort / flake8 / mypy clean; mkdocs build --strict clean.

Not done (left to the owner)

  • Merge to main, PyPI publish, and any launch post.

🤖 Generated with Claude Code

0-uddeshya-0 and others added 3 commits July 6, 2026 22:55
Map an agent skill's tests.json acceptance cases onto @behavior tests:
each case's input is handed to a user-supplied runner, expected.contains
becomes the must_contain guardrail, and the output is snapshotted so a
SKILL.md edit that keeps the required substrings but changes what the
skill says still fails CI as semantic drift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A german-b2b-outreach-email skill folder (SKILL.md + tests.json) plus a
deterministic canned runner show the full flow; the examples double as
dogfooding in CI's record/check step.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
README gains a 'Why BehaviorCI now' section (agent skills widen the
prompt-regression surface) and a worked from_skill_tests() example; the
docs site gains an 'Agent skills' page covering the tests.json mapping,
failure output, options, and runner hygiene. Changelog updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 7, 2026 05:55
@vouch-review

vouch-review Bot commented Jul 7, 2026 •

Copy link
Copy Markdown

Vouch Analysis

Analysis 4a06e27c-715d-4f77-8ef6-1592f6958508 · model deterministic analysis · confidence threshold 0.7

🛡️ Maintainer Gate — 2/4 evidence checks passed

Note

Some evidence checks did not pass (advisory for this author).

  • ❌ All packages exist on their registries — 7 packages could not be found on any registry — hallucinated or slopsquatting bait. Remove or replace before review.
  • ✅ No secrets or vulnerable versions introduced — No credential patterns or known-CVE versions in the diff.
  • ✅ Tests accompany code changes — 4 test files updated alongside 2 source files.
  • ⚠️ Dependency hygiene — 3 advisory dependency findings (unused/redundant packages). Advisory — never blocks on its own.

Warning

Security or hallucination findings were detected in this pull request.

🚨 Security & Hallucinations

Package/File Type Reason Details
module hallucination PyPI project not found: module View Details
follow_imports hallucination PyPI project not found: follow_imports View Details
follow_imports hallucination PyPI project not found: follow_imports View Details
follow_imports_for_stubs hallucination PyPI project not found: follow_imports_for_stubs View Details
from .api import behavior hallucination PyPI project not found: api View Details
from .exceptions import ConfigurationError hallucination PyPI project not found: exceptions View Details
from tests.test_plugin import MOCK_CONFTEST hallucination PyPI project not found: tests View Details

Note

Dependency quality suggestions are advisory and do not block merge on their own.

🧹 Dependency Quality

Package/File Type Reason Details
module anti-pattern Unused dependency: module View Details
follow_imports anti-pattern Unused dependency: follow_imports View Details
follow_imports_for_stubs anti-pattern Unused dependency: follow_imports_for_stubs View Details

🤖 AI System Disclosure: This analysis was generated by Vouch using deterministic AST parsing and registry verification (no LLM inference on this run). Learn more about how Vouch works.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds first-class “agent skill” support to BehaviorCI by introducing an adapter that converts a skill folder’s tests.json cases into standard @behavior snapshot tests, along with documentation and runnable examples to dogfood the record/check workflow in CI.

Changes:

  • Introduces from_skill_tests() adapter to generate one @behavior test per tests.json case (with expected.contains → must_contain mapping).
  • Adds direct + pytester end-to-end tests plus a runnable example skill fixture under tests/examples/.
  • Adds new docs page and updates README, MkDocs nav, and changelog to document skill testing.

Reviewed changes

Copilot reviewed 11 out of 11 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
tests/test_skills.py Adds unit + pytester end-to-end coverage for from_skill_tests() generation/validation and record→check behavior.
tests/examples/test_skill.py Provides a deterministic, runnable example showing how to wire from_skill_tests() into a test module.
tests/examples/skills/german-b2b-outreach-email/tests.json Adds example acceptance cases (input + expected.contains) for the sample skill.
tests/examples/skills/german-b2b-outreach-email/SKILL.md Adds example skill instructions used by the adapter example.
src/behaviorci/skills.py Implements the new skill adapter (from_skill_tests) and validation of tests.json.
src/behaviorci/init.py Exports from_skill_tests from the top-level package API.
README.md Documents “Testing agent skills” and positions BehaviorCI’s value for skills/drift.
mkdocs.yml Adds the new “Agent skills” page to the docs navigation.
docs/skills.md Adds dedicated documentation for skill folder format, adapter usage, and failure output.
docs/index.md Links the new Agent skills docs page from the docs landing page.
CHANGELOG.md Records the new adapter/docs/example under Unreleased additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/behaviorci/skills.py Outdated
skill_path = Path(skill_dir)
cases = _load_cases(skill_path / TESTS_FILENAME)

prefix = id_prefix or skill_path.resolve().name
0-uddeshya-0 and others added 2 commits July 7, 2026 12:52
numpy>=2.5 ships type aliases in its .pyi (valid only for target
>=3.12); under our 3.10 mypy target these are rejected as syntax errors,
failing the lint job on a fresh dependency resolve even though our own
code is unchanged. numpy is third-party (already under
--ignore-missing-imports), so skip following into its stubs. Verified
green on mypy 1.20.0 and 2.1.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Addresses review feedback: skill_path.resolve().name follows symlinks,
so a symlinked skill folder produced test ids from the link target's
name rather than the folder the caller passed. Use skill_path.name
(falling back to resolve() only for '.'/'/'), so ids are stable and
predictable across symlinked skill dirs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@0-uddeshya-0
0-uddeshya-0 merged commit 507fce6 into main Jul 7, 2026
5 of 6 checks passed
@0-uddeshya-0
0-uddeshya-0 deleted the revival branch July 7, 2026 11:19
0-uddeshya-0 added a commit that referenced this pull request Aug 25, 2026
Adds from_skill_tests() to turn a skill folder's tests.json into @behavior snapshot tests, with docs + runnable example. Includes CI fixes: numpy PEP695-stub mypy override, and symlink-stable test ids. All checks green.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants