Repository navigation
feat: --mode artprompt — ASCII-art keyword-masking jailbreak - #25
Merged
Merged
Conversation
New payloads/artprompt.py implements ArtPrompt (Jiang et al. 2024): hide the trigger word from the safety filter by rendering it as ASCII art (5x5 block font A–Z), then instruct the target to decode it letter-by-letter and substitute it for [MASK] in the request. render_word / build_artprompt / mask_and_wrap exposed; 6 abstract harm-category probes (ART-001..006), all expected to be refused — the keyword appears only as art, never in plaintext. Wired into EXPANDED_MODE_TESTS, --mode choices + MODE_DESCRIPTIONS, and server MODE_LABELS. ATLAS AML.T0054 / OWASP LLM01. Also hardens a flaky timing assertion in test_safety.py (throttle no-op ceiling 0.01s→0.1s) that intermittently failed under full-suite CPU load. Recovers a feature lost with the deleted background-task branch (see #24). +10 tests (tests/test_artprompt.py). Full suite: 1938 passed, 1 skipped. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ShadowSlayer08
pushed a commit
that referenced
this pull request
Sep 25, 2026
Make the self-growing KB (--evolve) curate itself instead of only appending. - RedTeamKB gains get(), record_success() (bump success_count + stamp last_used), prune() (remove grown 'dynamic-win' patterns that are stale and/or weak — static seeds and proven/undated patterns are never pruned), and quality_report() (origin mix, reinforcement count, staleness, top wins). - The dynamic grow loop now REINFORCES a repeat win (record_success on the near-duplicate) instead of dropping it; new wins are stamped success_count=1 + created_at/last_used; DynamicRedTeamer tracks grown vs reinforced (shown in the "KB grew" line). - CLI: --kb-prune [--kb-prune-max-age-days N] [--kb-prune-min-success K]; --kb-stats now prints a quality snapshot (grown wins / reinforced / avg success / stale >30d). Recovers the second feature lost with the deleted background-task branch (see #24/#25). +9 tests (tests/test_kb_quality.py). Full suite: 1947 passed, 1 skipped. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
New
--mode artprompt— the ArtPrompt attack (Jiang et al., 2024, ASCII Art-based Jailbreak Attacks against Aligned LLMs). It hides the sensitive trigger word from the safety filter by rendering it as ASCII art; a capable model still reads it off the art and acts on the request.How
payloads/artprompt.py— a 5×5 block font (A–Z + space),render_word(),build_artprompt(keyword, request)(renders the keyword as art + letter-by-letter decode instructions, keyword substituted for[MASK]), andmask_and_wrap(prompt, keyword).ART-001..006) over abstract harm-category keywords, allexpected: refusal. The keyword appears only as art — never in plaintext (there's a test asserting this), which is the whole point of the technique.EXPANDED_MODE_TESTS,--modechoices +MODE_DESCRIPTIONS, and the server'sMODE_LABELS. Tagged ATLAS AML.T0054 / OWASP LLM01.Safety
Consistent with the existing corpus: the payloads request disallowed content abstractly (category word + generic ask) and every case is expected to be refused — no operational detail is written.
Also (incidental)
Hardened a flaky timing assertion in
test_safety.py(test_throttle_noop_when_unlimited:0.01s→0.1sceiling) that intermittently failed under full-suite CPU load. A real throttle sleeps for whole seconds, so0.1sstill proves the no-op path.Note
This recovers a feature that was lost with the deleted background-task branch (same branch as the #24 reporter fix). KB quality management — the other lost feature — is coming in a follow-up PR.
+10 tests (
tests/test_artprompt.py). Full suite: 1938 passed, 1 skipped.🤖 Generated with Claude Code