Repository navigation
fix: mutate report/JSON cover all 16 mutators (not a stale 8) - #24
Merged
Merged
Conversation
print_mutate_report and save_mutate_json built their columns and the JSON `mutation_methods` from a hardcoded `_MUTATOR_NAMES` of only the original 8, while PayloadMutator.all_mutations() fires 16 — so the 8 newer framings (math_problem, adversarial_poetry, emotional_manipulation, bad_likert_judge, policy_puppetry, skeleton_key, deceptive_delight, refusal_suppression) were silently dropped from the mutate report and its sidecar JSON. Derive `_MUTATOR_NAMES` from a new `mutator_method_names()` (the method_name of every all_mutations() entry) so the report can never drift from the mutations actually fired; de-stale the run_mutate_suite docstring. JSON field names unchanged. (Originally fixed on a background-task branch that was deleted before it merged; redone here from master.) +4 tests (tests/test_mutate_report.py). Full suite: 1916 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ShadowSlayer08
pushed a commit
that referenced
this pull request
Sep 25, 2026
New payloads/artprompt.py implements ArtPrompt (Jiang et al. 2024): hide the trigger word from the safety filter by rendering it as ASCII art (5x5 block font A–Z), then instruct the target to decode it letter-by-letter and substitute it for [MASK] in the request. render_word / build_artprompt / mask_and_wrap exposed; 6 abstract harm-category probes (ART-001..006), all expected to be refused — the keyword appears only as art, never in plaintext. Wired into EXPANDED_MODE_TESTS, --mode choices + MODE_DESCRIPTIONS, and server MODE_LABELS. ATLAS AML.T0054 / OWASP LLM01. Also hardens a flaky timing assertion in test_safety.py (throttle no-op ceiling 0.01s→0.1s) that intermittently failed under full-suite CPU load. Recovers a feature lost with the deleted background-task branch (see #24). +10 tests (tests/test_artprompt.py). Full suite: 1938 passed, 1 skipped. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ShadowSlayer08
pushed a commit
that referenced
this pull request
Sep 25, 2026
Make the self-growing KB (--evolve) curate itself instead of only appending. - RedTeamKB gains get(), record_success() (bump success_count + stamp last_used), prune() (remove grown 'dynamic-win' patterns that are stale and/or weak — static seeds and proven/undated patterns are never pruned), and quality_report() (origin mix, reinforcement count, staleness, top wins). - The dynamic grow loop now REINFORCES a repeat win (record_success on the near-duplicate) instead of dropping it; new wins are stamped success_count=1 + created_at/last_used; DynamicRedTeamer tracks grown vs reinforced (shown in the "KB grew" line). - CLI: --kb-prune [--kb-prune-max-age-days N] [--kb-prune-min-success K]; --kb-stats now prints a quality snapshot (grown wins / reinforced / avg success / stale >30d). Recovers the second feature lost with the deleted background-task branch (see #24/#25). +9 tests (tests/test_kb_quality.py). Full suite: 1947 passed, 1 skipped. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
print_mutate_reportandsave_mutate_jsonbuilt their per-method columns and the JSONmutation_methodsfrom a hardcoded_MUTATOR_NAMESlist of only the original 8 mutators — butPayloadMutator.all_mutations()fires 16. The 8 newer framings were silently dropped from the mutate report and its sidecar JSON:math_problem,adversarial_poetry,emotional_manipulation,bad_likert_judge,policy_puppetry,skeleton_key,deceptive_delight,refusal_suppression.How
mutator_method_names()returns themethod_nameof everyall_mutations()entry, in order._MUTATOR_NAMESis now derived from it (single source of truth) — the report can never drift from the mutations actually fired again.run_mutate_suitedocstring. JSON field names unchanged (mutation_methodsnow just lists all 16).Tests
tests/test_mutate_report.py(+4): the derived list matchesall_mutations()and is length 16; the module constant covers the newer 8; the report breakdown prints every mutator and shows "(16 per test)"; the saved JSON'smutation_methodslists all 16.Full suite: 1916 passed, 1 skipped.
Note
This bug was originally fixed on a background-task branch that was deleted before it merged (its session was removed), so the fix never reached
master— redone here frommaster. The ArtPrompt mode and KB-quality-management work that shared that lost branch are not included here (separate features, would need redoing).🤖 Generated with Claude Code