Skip to content

feat: --mode artprompt — ASCII-art keyword-masking jailbreak - #25

Merged
ShadowSlayer08 merged 1 commit into
masterfrom
feat/artprompt-mode
Sep 25, 2026
Merged

ShadowSlayer08 merged 1 commit into
masterfrom
feat/artprompt-mode

Conversation

@ShadowSlayer08

Copy link
Copy Markdown
Owner

What

New --mode artprompt — the ArtPrompt attack (Jiang et al., 2024, ASCII Art-based Jailbreak Attacks against Aligned LLMs). It hides the sensitive trigger word from the safety filter by rendering it as ASCII art; a capable model still reads it off the art and acts on the request.

crucible --mode artprompt --local

How

  • payloads/artprompt.py — a 5×5 block font (A–Z + space), render_word(), build_artprompt(keyword, request) (renders the keyword as art + letter-by-letter decode instructions, keyword substituted for [MASK]), and mask_and_wrap(prompt, keyword).
  • 6 probes (ART-001..006) over abstract harm-category keywords, all expected: refusal. The keyword appears only as art — never in plaintext (there's a test asserting this), which is the whole point of the technique.
  • Wired into EXPANDED_MODE_TESTS, --mode choices + MODE_DESCRIPTIONS, and the server's MODE_LABELS. Tagged ATLAS AML.T0054 / OWASP LLM01.

Safety

Consistent with the existing corpus: the payloads request disallowed content abstractly (category word + generic ask) and every case is expected to be refused — no operational detail is written.

Also (incidental)

Hardened a flaky timing assertion in test_safety.py (test_throttle_noop_when_unlimited: 0.01s → 0.1s ceiling) that intermittently failed under full-suite CPU load. A real throttle sleeps for whole seconds, so 0.1s still proves the no-op path.

Note

This recovers a feature that was lost with the deleted background-task branch (same branch as the #24 reporter fix). KB quality management — the other lost feature — is coming in a follow-up PR.

+10 tests (tests/test_artprompt.py). Full suite: 1938 passed, 1 skipped.

🤖 Generated with Claude Code

New payloads/artprompt.py implements ArtPrompt (Jiang et al. 2024): hide the trigger
word from the safety filter by rendering it as ASCII art (5x5 block font A–Z), then
instruct the target to decode it letter-by-letter and substitute it for [MASK] in the
request. render_word / build_artprompt / mask_and_wrap exposed; 6 abstract
harm-category probes (ART-001..006), all expected to be refused — the keyword appears
only as art, never in plaintext. Wired into EXPANDED_MODE_TESTS, --mode choices +
MODE_DESCRIPTIONS, and server MODE_LABELS. ATLAS AML.T0054 / OWASP LLM01.

Also hardens a flaky timing assertion in test_safety.py (throttle no-op ceiling
0.01s→0.1s) that intermittently failed under full-suite CPU load.

Recovers a feature lost with the deleted background-task branch (see #24).
+10 tests (tests/test_artprompt.py). Full suite: 1938 passed, 1 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@ShadowSlayer08
ShadowSlayer08 merged commit 6902017 into master Sep 25, 2026
6 checks passed
@ShadowSlayer08
ShadowSlayer08 deleted the feat/artprompt-mode branch September 25, 2026 01:33
ShadowSlayer08 pushed a commit that referenced this pull request Sep 25, 2026
Make the self-growing KB (--evolve) curate itself instead of only appending.

- RedTeamKB gains get(), record_success() (bump success_count + stamp last_used),
  prune() (remove grown 'dynamic-win' patterns that are stale and/or weak — static
  seeds and proven/undated patterns are never pruned), and quality_report() (origin
  mix, reinforcement count, staleness, top wins).
- The dynamic grow loop now REINFORCES a repeat win (record_success on the
  near-duplicate) instead of dropping it; new wins are stamped success_count=1 +
  created_at/last_used; DynamicRedTeamer tracks grown vs reinforced (shown in the
  "KB grew" line).
- CLI: --kb-prune [--kb-prune-max-age-days N] [--kb-prune-min-success K]; --kb-stats
  now prints a quality snapshot (grown wins / reinforced / avg success / stale >30d).

Recovers the second feature lost with the deleted background-task branch (see #24/#25).
+9 tests (tests/test_kb_quality.py). Full suite: 1947 passed, 1 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant