Find sensitive information in cloned repositories to build personalized guardrail lists, so that sensitive data is not leaked.
Written in Go — a single static binary with both an interactive TUI and a headless CLI.
Two scanning approaches:
- Regex scanner — fast, pattern-based detection across files and git history (45 built-in patterns)
- Ollama scanner — LLM-powered, context-aware detection using a local Ollama model
Both support scanning a single repository or a parent directory containing hundreds of cloned repos, and every run is persisted to a timestamped results directory with a README listing the scanned paths and their git SHAs.
flowchart TB
TUI["Interactive TUI<br/>llm-redact"]
CLI["Headless CLI<br/>llm-redact scan / ollama"]
CFG["Scan configuration<br/>path · scanners · globs · workers"]
DISC["Repository discovery<br/>walks the path for .git dirs, up to --depth"]
TUI --> CFG
CLI --> CFG
CFG --> DISC
DISC --> ENGINE
subgraph ENGINE["Scan engine — repos in parallel (--processes), files in parallel (--workers)"]
direction TB
FILES["Working-tree files<br/>include/exclude globs · binary files skipped"]
HIST["Git history<br/>commit messages + added diff lines<br/>(--history-depth)"]
RED["Suppression pass<br/>every regex.txt match is replaced<br/>with a REDACTED placeholder"]
CHOICE{"scanner?"}
RGX["Regex scanner<br/>45 patterns · 8 categories<br/>Luhn + placeholder filters"]
OLL["Ollama scanner<br/>streaming /api/generate<br/>JSON findings from the model"]
FILES --> RED
HIST --> RED
RED --> CHOICE
CHOICE -->|regex| RGX
CHOICE -->|ollama| OLL
end
ENGINE -. "progress events<br/>(files, commits, findings, model tokens)" .-> LIVE["Live progress<br/>TUI screens / CLI stderr"]
RGX --> FIND["Findings<br/>severity · category · file:line · commit"]
OLL --> FIND
FIND --> TERM["Terminal output<br/>table · json · txt · csv"]
FIND --> RESDIR["results/<date_time>/<br/>README.md — paths + git SHAs<br/>findings.json · config.json · extras"]
The engine is UI-agnostic: the TUI and the CLI consume the same progress-event stream, and both persist every run to the results directory — including partial results when a scan is cancelled.
- Go 1.22+ (to build) — or grab a prebuilt binary
- Git (for commit history scanning and HEAD SHAs)
- Ollama (only for the LLM-based scanner)
# 1. Build
make build # or: go build -o llm-redact ./cmd/llm-redact
# 2. Open the interactive TUI
./llm-redact
# 3. …or scan headlessly with regex patterns
./llm-redact scan --path /path/to/repos
# 4. …or scan with a local Ollama model
./llm-redact ollama --path /path/to/repos --model llama3Install straight from the repo instead:
go install github.com/lab34-es/llm-redact/cmd/llm-redact@latestRun llm-redact with no arguments (or llm-redact tui) in a terminal:
-
Path selection — type the path to scan and the discovery depth.
-
Repository selection — every discovered repo is listed with its HEAD SHA; toggle with
space,aselects all,nnone. -
Configuration — choose the scanner (
←/→). For Ollama the model picker is filled from the server's model list (/api/tags); if the server is unreachable you can type a model name. Set include/exclude globs, a suppression regex file, and git-history options. -
Live progress — overall repo progress bar, per-file progress, a live findings ticker, and (for Ollama) the model's streaming output for the file being analyzed.
qcancels; partial results are still saved. -
Results browser — findings table with severity filters (
h/m/l,afor all),enteropens a detail view,nstarts a new scan. The footer shows where results were saved.
llm-redact tui --path /repos --depth 3 prefills the first screen.
Scans files and git history (commit messages and diffs) using 45 compiled regex patterns across 8 categories:
- API keys and tokens (AWS, GitHub, GitLab, Slack, Stripe, Google, Heroku, Twilio, SendGrid, npm, PyPI, JWTs, generic)
- Passwords and secrets (assignments, env vars, basic auth URLs)
- Private keys and certificates (RSA, DSA, EC, OpenSSH, PGP, PKCS8)
- Database connection strings (PostgreSQL, MySQL, MongoDB, Redis, MSSQL, JDBC)
- IP addresses and internal URLs (private ranges, localhost)
- Email addresses
- Credit card numbers (Visa, Mastercard, Amex, Discover — Luhn-validated)
- SSN and national IDs (US SSN, UK NINO)
# Scan a single repo, print a table
./llm-redact scan --path /path/to/repo
# Scan a directory with many cloned repos
./llm-redact scan --path /path/to/repos --depth 2
# Persist per-repo JSON / TXT, or a single CSV
./llm-redact scan --path /path/to/repos --output json
./llm-redact scan --path /path/to/repos --output csv --output-dir ./reports
# Only scan Python and YAML files, skip git history
./llm-redact scan --path /path/to/repos --include '*.py' --include '*.yml' --no-git-history
# Limit git history scanning to the last 100 commits per repo
./llm-redact scan --path /path/to/repos --history-depth 100
# 8 parallel workers per repo, 4 repos in parallel
./llm-redact scan --path /path/to/repos --workers 8 --processes 4| Flag | Description | Default |
|---|---|---|
--path |
Path to a repo or directory containing repos | required |
--depth |
Max depth to search for git repos | 2 |
--output |
Output format: table, json, txt, csv |
table |
--output-dir |
Base directory for the timestamped results dir | ./results |
--include |
Glob for files to include (repeat or comma-separate) | all files |
--exclude |
Glob for files to exclude (repeat or comma-separate) | none |
--regex-file |
Suppression patterns (see below) | - |
--no-git-history |
Skip commit message and diff scanning | false |
--history-depth |
Max commits to scan per repo (0 = all) | all |
--workers |
Parallel workers for file scanning | CPU count |
--processes |
Repos scanned in parallel | 1 |
--quiet |
No per-finding progress lines on stderr | false |
--no-color |
Disable colored output (NO_COLOR is honored too) | false |
Migration note:
--include '*.py' '*.yml'(space-separated values) from the Python version is now--include '*.py' --include '*.yml'or--include '*.py,*.yml'. A stray positional argument is a hard usage error so the change cannot pass silently.
Sends each file to a local Ollama model for context-aware sensitive-data detection. Catches things regex patterns miss (hardcoded internal hostnames, encoded secrets, sensitive comments).
ollama pull llama3 # prerequisite: a running Ollama with a model
./llm-redact ollama --path /path/to/repo --model llama3
./llm-redact ollama --path /path/to/repos --model mistral --depth 2
./llm-redact ollama --path /path/to/repos --model llama3 --output json
./llm-redact ollama --path /path/to/repos --model llama3 \
--include '*.yml' --include '*.json' --include '*.env*' --max-file-size 204800
./llm-redact ollama --path /path/to/repos --model llama3 --ollama-host http://192.168.1.100:11434| Flag | Description | Default |
|---|---|---|
--model |
Ollama model name | required |
--max-file-size |
Skip files larger than N bytes | 102400 (100KB) |
--ollama-host |
Ollama API URL | http://localhost:11434 |
--ollama-timeout |
Seconds without model output before a file is aborted | 120 |
Plus all the shared flags above (--path, --depth, --output, --include, …).
Every run — CLI or TUI — is persisted to <output-dir>/<date_time>/ (default ./results/):
results/2026-08-22_14-05-33/
├── README.md # scanner, config, and a table of scanned paths with their git SHAs
├── findings.json # full machine-readable results (config + all findings)
├── config.json # the exact resolved configuration, for reproducibility
└── … # extras per --output: <repo>.json / <repo>.txt per repo, or findings.csv
The README's path table looks like:
| Path | Git SHA | Findings | Files scanned | Commits scanned |
|---|---|---|---|---|
/repos/org-a/api |
3f2a1c9… |
12 | 240 | 87 |
Interrupted runs (ctrl+c, or q in the TUI) still write the directory with the partial results and a note.
--output chooses the terminal format and which extra files appear; --output-dir moves the base directory.
| Code | Meaning |
|---|---|
0 |
Scan completed, no findings |
1 |
Scan completed, findings present (useful in CI) |
2 |
Usage error (bad flags) |
3 |
Runtime failure (unreachable Ollama, unreadable path, …) |
--regex-file points at a file with one Go (RE2) regular expression per line (# comments allowed). Every match is replaced with [REDACTED] before scanning — in files and in git history — suppressing findings you have allowlisted. See the annotated regex.txt.
RE2 has no lookahead/lookbehind/backreferences; patterns written for Python's
remay need adjusting. Invalid patterns fail fast with their line number.
When --path points to a directory that is not itself a git repo, the tool searches for .git/ directories up to --depth levels deep:
/repos/
org-a/
api-service/ <- git repo
frontend/ <- git repo
org-b/
backend/ <- git repo
./llm-redact scan --path /repos --depth 2 discovers all three. If nothing is found, the path itself is scanned as a plain directory (no git SHA in the results README).
llm-redact/
├── cmd/llm-redact/ # entry point
├── internal/
│ ├── model/ # domain types + progress events
│ ├── patterns/ # 45 compiled regex patterns across 8 categories
│ ├── redact/ # regex.txt suppression ([REDACTED])
│ ├── discovery/ # repo discovery, file walking, binary detection
│ ├── gitscan/ # git history scanning (messages + diffs) via git CLI
│ ├── ollama/ # minimal streaming Ollama client
│ ├── engine/ # scan orchestration + worker pools + events
│ ├── output/ # table / json / txt / csv formatters
│ ├── results/ # results/<date_time>/ writer
│ ├── cli/ # flags, subcommands, headless progress
│ └── tui/ # Bubble Tea interactive UI
├── regex.txt # example suppression list
├── Makefile
└── README.md
make build # build ./llm-redact
make test # go test ./...
make vet # go vet ./...The test suite runs fully offline: the Ollama client is tested against local fake HTTP servers, and git tests build fixture repositories on the fly.




