GPU-backed HTTP wrapper around two PII models plus a co-hosted Gemma chat model, in one Tinfoil container:
screenpipe/pii-redactor(v50_distilled6l) — six-layer, vocab-pruned XLM-R token classifier for text PII. EndpointPOST /filter.screenpipe/pii-image-redactor(rfdetr_v19) — RF-DETR-Nano detector for image PII in screenshots, trained on real-app and synthetic screens. EndpointPOST /image/detect.google/gemma-4-E2B-it— chat + vision + audio via vLLM. EndpointPOST /v1/chat/completionswith the backwards-compatible API model id"gemma4-e4b". Weights are baked into the image (~10 GB BF16).
Note: the 31B model (
gemma4-31b) is not in this container. It stays on Tinfoil's hostedinference.tinfoil.shendpoint, reached separately bypackages/ai-gatewayin the Cloudflare worker. The co-hosted E2B is a smaller multimodal-audio model, not a drop-in replacement.
All three workloads deploy inside the same Tinfoil confidential-compute container on one H200 (~141 GB VRAM) so neither pixels, text, nor chat prompts leave an attested runtime. The shim only publishes uvicorn on :8080; /v1/* requests are reverse-proxied to vLLM on 127.0.0.1:8001. One TLS-attested URL, one auth token, three workloads sharing one allocation.
GET /health → {"status": "ok", "model_ready": true, "image_model_ready": true, ...}
POST /filter → {"text": "My email is alice@foo.com"}
← {"redacted": "My email is [EMAIL]",
"spans": [{"label": "private_email", "start": 12, "end": 25,
"text": "alice@foo.com", "score": 0.99}],
"latency_ms": 180,
"model": "screenpipe/pii-redactor:v50_distilled6l (mixed-int4-int8-onnx)"}
POST /image/detect → {"image_b64": "<b64-jpg-or-png>", "threshold": 0.30}
← {"detections": [{"bbox": [x, y, w, h], "label": "private_person", "score": 0.95},
{"bbox": [x, y, w, h], "label": "secret", "score": 0.91}],
"latency_ms": 32, "model": "rfdetr_v19",
"width": 2880, "height": 1800}
Bbox is [x, y, w, h] in original-image pixel space (the server un-resizes from the model's 512×512 input). Labels are the canonical 12-class screenpipe PII taxonomy.
# build
docker build -t privacy-filter:dev .
# run
docker run --rm -p 8080:8080 privacy-filter:dev
# smoke test
curl -s http://localhost:8080/health
curl -s -X POST http://localhost:8080/filter \
-H 'Content-Type: application/json' \
-d '{"text":"Call Alice at +1 415 555 0100 about alice@example.com"}' | jq
# Image: send a JPG/PNG as base64.
B64=$(base64 -i some_screenshot.png)
curl -s -X POST http://localhost:8080/image/detect \
-H 'Content-Type: application/json' \
-d "$(jq -nc --arg img "$B64" '{image_b64: $img, threshold: 0.30}')" | jqThe first build downloads the v50 text contract (~133 MB across model, tokenizer, config, and vocabulary remap), the 60 MB rfdetr_v19 ONNX, and the baked Gemma weights. Subsequent builds use the build cache. Both HuggingFace revisions and every model artifact checksum are pinned, so a rebuild cannot silently drift to different weights.
-
Build and push the image to GitHub Container Registry using the
privacy-filter releaseworkflow. Supply the semantic version without thevprefix (for example0.8.0). The workflow prints the immutable image digest.gh workflow run release.yml --ref <branch> -f version=0.8.0
-
Pin the digest in
tinfoil-config.yml, commit it, then create the matchingv*tag. The tag runs the Tinfoil measurement/attestation job.git add tinfoil-config.yml git commit -m "release: privacy-filter v0.8.0" git tag v0.8.0 && git push origin <branch> v0.8.0
-
Click-through in the Tinfoil dashboard (https://dash.tinfoil.sh):
- create an org (if you haven't already)
- connect the GitHub app to this repo
- pick the tag to deploy
- wait for status =
Running(cold start ~30–60 s for the first model load)
-
Verify — the service is now reachable at
https://privacy-filter.<org>.containers.tinfoil.dev/health.
Text model (v50_distilled6l) — mixed int4/int8 ONNX:
| Metric | Value |
|---|---|
| Model artifact | 114 MB |
| Model shape | six-layer XLM-R student |
| Runtime window | 256 tokens with 64-token overlap |
| Local CPU reference | 7–9 ms for short captured strings |
Image model (rfdetr_v19) — TensorRT/CUDA EP:
| Metric | Value |
|---|---|
| Weights (FP16, FP32 I/O) | 60 MB |
| Params | ~25 M |
| Input resolution | 512×512 |
| Local CPU reference | ~118 ms/frame |
| Real-app planted-secret gate | 19/19 detected at the production 0.50 floor |
The deployed CVM also carries Gemma E2B and its vLLM working set, so
tinfoil-config.yml remains the source of truth for total CPU, memory, and GPU sizing.
- Tinfoil remote attestation covers the exact image digest, so clients can verify the specific model bits + server code that handled their request.
- Only the paths listed in
tinfoil-config.ymlare exposed — Tinfoil's shim allowlist blocks every other URL at the enclave boundary. - Model weights are baked in. Text model: build-time HF download then
TRANSFORMERS_OFFLINE=1. Image model: build-time HF download with SHA-256 verification via Docker'sADD --checksum=. No runtime HuggingFace calls, so an attacker who subverts DNS can't swap weights out from under the enclave. - The exact container image is measured and attested. The process currently runs as root because the co-hosted vLLM needs access to NVIDIA device nodes; the Tinfoil shim and CVM remain the external security boundary.
- Text: English-primary. Multilingual coverage per upstream model card varies.
- Text: long inputs are split into overlapping 256-token windows; request size is capped separately by
MAX_INPUT_CHARS. - Image: v19 is validated on released eval suites and planted real-app secrets, but unusual layouts, scripts, handwriting, occlusion, and adversarial inputs still need their own evaluation.
- Image: payload cap
MAX_IMAGE_BYTES=20 MB(override via env). Decoded RGB working buffer is ~3× the payload. - Not a compliance certification. One layer in a privacy-by-design stack.
PolyForm Noncommercial License 1.0.0. You can read, run, fork, modify, and share — noncommercial use only. Commercial use requires a separate license; reach out to louis@screenpi.pe.
Bundled / downloaded model weights keep their own licenses: screenpipe/pii-redactor, screenpipe/pii-image-redactor (rfdetr_v19), and google/gemma-4-E2B-it (Gemma Terms of Use). Using this image commercially means complying with all of them in addition to this repo's license.