Skip to content

fix: mutate report/JSON cover all 16 mutators (not a stale 8) - #24

Merged
ShadowSlayer08 merged 1 commit into
masterfrom
fix/mutate-report-columns
Sep 25, 2026
Merged

ShadowSlayer08 merged 1 commit into
masterfrom
fix/mutate-report-columns

Conversation

@ShadowSlayer08

Copy link
Copy Markdown
Owner

What

print_mutate_report and save_mutate_json built their per-method columns and the JSON mutation_methods from a hardcoded _MUTATOR_NAMES list of only the original 8 mutators — but PayloadMutator.all_mutations() fires 16. The 8 newer framings were silently dropped from the mutate report and its sidecar JSON:

math_problem, adversarial_poetry, emotional_manipulation, bad_likert_judge, policy_puppetry, skeleton_key, deceptive_delight, refusal_suppression.

How

  • New mutator_method_names() returns the method_name of every all_mutations() entry, in order.
  • _MUTATOR_NAMES is now derived from it (single source of truth) — the report can never drift from the mutations actually fired again.
  • De-staled the run_mutate_suite docstring. JSON field names unchanged (mutation_methods now just lists all 16).

Tests

tests/test_mutate_report.py (+4): the derived list matches all_mutations() and is length 16; the module constant covers the newer 8; the report breakdown prints every mutator and shows "(16 per test)"; the saved JSON's mutation_methods lists all 16.

Full suite: 1916 passed, 1 skipped.

Note

This bug was originally fixed on a background-task branch that was deleted before it merged (its session was removed), so the fix never reached master — redone here from master. The ArtPrompt mode and KB-quality-management work that shared that lost branch are not included here (separate features, would need redoing).

🤖 Generated with Claude Code

print_mutate_report and save_mutate_json built their columns and the JSON
`mutation_methods` from a hardcoded `_MUTATOR_NAMES` of only the original 8, while
PayloadMutator.all_mutations() fires 16 — so the 8 newer framings (math_problem,
adversarial_poetry, emotional_manipulation, bad_likert_judge, policy_puppetry,
skeleton_key, deceptive_delight, refusal_suppression) were silently dropped from the
mutate report and its sidecar JSON.

Derive `_MUTATOR_NAMES` from a new `mutator_method_names()` (the method_name of every
all_mutations() entry) so the report can never drift from the mutations actually
fired; de-stale the run_mutate_suite docstring. JSON field names unchanged.

(Originally fixed on a background-task branch that was deleted before it merged;
redone here from master.) +4 tests (tests/test_mutate_report.py). Full suite: 1916 passed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@ShadowSlayer08
ShadowSlayer08 merged commit 5951e90 into master Sep 25, 2026
6 checks passed
@ShadowSlayer08
ShadowSlayer08 deleted the fix/mutate-report-columns branch September 25, 2026 01:17
ShadowSlayer08 pushed a commit that referenced this pull request Sep 25, 2026
New payloads/artprompt.py implements ArtPrompt (Jiang et al. 2024): hide the trigger
word from the safety filter by rendering it as ASCII art (5x5 block font A–Z), then
instruct the target to decode it letter-by-letter and substitute it for [MASK] in the
request. render_word / build_artprompt / mask_and_wrap exposed; 6 abstract
harm-category probes (ART-001..006), all expected to be refused — the keyword appears
only as art, never in plaintext. Wired into EXPANDED_MODE_TESTS, --mode choices +
MODE_DESCRIPTIONS, and server MODE_LABELS. ATLAS AML.T0054 / OWASP LLM01.

Also hardens a flaky timing assertion in test_safety.py (throttle no-op ceiling
0.01s→0.1s) that intermittently failed under full-suite CPU load.

Recovers a feature lost with the deleted background-task branch (see #24).
+10 tests (tests/test_artprompt.py). Full suite: 1938 passed, 1 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ShadowSlayer08 pushed a commit that referenced this pull request Sep 25, 2026
Make the self-growing KB (--evolve) curate itself instead of only appending.

- RedTeamKB gains get(), record_success() (bump success_count + stamp last_used),
  prune() (remove grown 'dynamic-win' patterns that are stale and/or weak — static
  seeds and proven/undated patterns are never pruned), and quality_report() (origin
  mix, reinforcement count, staleness, top wins).
- The dynamic grow loop now REINFORCES a repeat win (record_success on the
  near-duplicate) instead of dropping it; new wins are stamped success_count=1 +
  created_at/last_used; DynamicRedTeamer tracks grown vs reinforced (shown in the
  "KB grew" line).
- CLI: --kb-prune [--kb-prune-max-age-days N] [--kb-prune-min-success K]; --kb-stats
  now prints a quality snapshot (grown wins / reinforced / avg success / stale >30d).

Recovers the second feature lost with the deleted background-task branch (see #24/#25).
+9 tests (tests/test_kb_quality.py). Full suite: 1947 passed, 1 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant