Skip to content

feat(opf): scan in the background so git push never waits on the privacy filter - #2631

Open
peyton-alt wants to merge 23 commits into
peyton/opf-batch-cap-fixfrom
peyton/opf-scan-worker
Open

peyton-alt wants to merge 23 commits into
peyton/opf-batch-cap-fixfrom
peyton/opf-scan-worker

Conversation

@peyton-alt

@peyton-alt peyton-alt commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

https://entire.io/gh/entireio/cli/trails/1457

Stacked on #2533, which merges first; this PR depends on it.

Problem

OPF inference is slow, about 1.14 s per KB of prose on CPU, so one real session is minutes to hours of model time. The model ran inside the pre-push hook, which left two bad options. With the old 2 MiB cap, every real session was refused, so OPF could not ship real checkpoints at all. With the cap raised, git push would wait for the whole scan, which on git-branch (the default backend) could take hours.

Result

git push never waits on the model, on either backend. Checkpoints still reach the remote only after OPF has scanned them.

  • Scan and apply are split (redact). ScanBlobsWithPrivacyFilter runs the model and stores results in a per-blob span cache. ApplyCachedPrivacyFilter redacts from the cache alone and reports anything not scanned yet. The cache lives in entire-opf-cache/ in the git common directory. It stores only SHA-256 hashes of leaf text plus the offsets and labels OPF found, never text, keyed by blob hash and category set. Entries older than 30 days are pruned.

  • Pre-push rewrites from the cache. Anything not scanned yet is held back: on git-refs that ref stays queued while scanned siblings ship, and on git-branch v1 is not pushed this time. The user sees [entire] Checkpoints are being scanned by the OpenAI Privacy Filter in the background and will be pushed when it finishes. and their own push completes.

  • Background worker (entire __opf_scan <remote>, one per repo via a lock):

    • collects the untrailered checkpoint commits on the primary backend
    • scans whatever the cache lacks
    • runs the same pre-push delivery for the remote the push named, using that push's OPF decision and non-interactive git

    If it cannot deliver, the checkpoints stay queued and the next push sends them from the cache without calling the model.

  • Capacity: model calls are split into chunks of at most 1 MiB, each with a deadline scaled to its size (30s floor, 3h ceiling). The caps are backstops for pathological input again: 128 MiB prose, 256 MiB raw.

  • Visibility: entire status (text and --json checkpoint_opf_pending) counts held checkpoints on both backends. git-branch users with held checkpoints get a tip pointing at entire doctor migrate-checkpoints, at most once every 30 days.

  • Unchanged: the explicit migration push still scans synchronously. No categories enabled, divergence and a Ctrl-C at the prompt behave as before.

  • Carried over from the earlier version of fix(strategy): dedup OPF byte cap, scope it per checkpoint ref, run it off git push's blocking path #2533: SpawnDetached gives children the null device instead of a pipe. With a pipe, the child was killed on its first stderr write once git push returned.

Where to start reviewing

  1. redact/opf_cache.go and the redact/batch.go refactor
  2. strategy/manual_commit_opf_scan.go: the worker, the spawn decision and the lock
  3. strategy/manual_commit_push.go: pre-push gating on both backends
  4. strategy/manual_commit_opf_rewrite.go / _refs.go: the rewrite modes

Known gaps (not in this PR)

  • An explicit git push origin entire/checkpoints/v1 is not blocked while v1 is unscanned. Pre-push does not read the refs git is pushing, and a rewrite inside the hook cannot change what git sends. This is the same on main.
  • In CI or a throwaway container, the worker usually dies with the job, so those checkpoints stay local rather than shipping unscanned.
  • entire clean does not reclaim entire-opf-cache/ yet. The worker's 30-day pruning bounds it.

Validation

  • mise run check: lint, unit and integration tests, Vogon 56/56, Roger-Roger 4/4.
  • New tests:
    • cache round trip, invalid keys and pruning
    • cache resolved from the repo, not the working directory
    • scan then apply gives the same output as the one-pass batch
    • pending, category-change and partial-entry cases; the breaker; runtime failure storing nothing
    • cache-only mode on both backends
    • pre-push holds and spawns on both backends
    • the worker delivers on both backends; lock exclusion; runtime failure delivers nothing
    • pending counts; hint shown once
    • an unusable cache is reported with its cause, not as a scan in progress
    • FuzzApplyCachedPrivacyFilter: for arbitrary blob contents, cached apply equals the one-pass batch, and a missing leaf makes it report pending. Fuzzed for 90s (~610k inputs) with no failures.
  • The OPF command-trust integration tests now wait for the worker to finish before checking whether the payload ran.
  • Real binary with the real opf model and a local bare remote:
    • git-refs, two checkpoints: git push took 0s with both held, and the worker delivered both about 20s later without another push. On the remote they are trailered, with 0 cleartext names and [REDACTED_PERSON] in their place.
    • git-branch: git push took 1s with v1 held, and the worker pushed v1 about 15s later, redacted.
    • No categories enabled: withheld with the config error, no worker started.
    • Remote rejects checkpoint refs: the worker rewrote locally but delivery was refused and the ref stayed queued. The next push delivered it in 1s from the cache, with no model call.
    • Second push during a ~45s scan: returned immediately, no second worker. The running worker picked up the new checkpoint and delivered both.

🤖 Generated with Claude Code

peyton-alt and others added 7 commits September 30, 2026 10:38
BatchBytesWithPrivacyFilter ran the model and applied its spans in one call,
so whoever needed redacted output had to wait for inference. Split it:

- ScanBlobsWithPrivacyFilter runs OPF over blobs with no cache entry and
  stores, per blob, the spans for each prose leaf, keyed by the blob's object
  hash and the category set. Entries hold leaf hashes and offsets, no text.
- ApplyCachedPrivacyFilter redacts from the cache alone and returns
  ErrOPFScanPending unless every blob and every leaf is covered.
- BatchBytesWithPrivacyFilter keeps its behavior and shares the same leaf
  collection and scan code.

Model calls are split into chunks of at most 1 MiB of leaf text, each with a
deadline scaled to its size (floor 30s, ceiling 3h) instead of a fixed 30s,
and the batch input ceiling rises to 256 MiB. The throughput and bug
investigation docs these numbers come from are included.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3SPBMJ9H9KF399NA898R42W
Implements redact.OPFSpanCache as one small JSON file per blob under
entire-opf-cache/, written atomically through the shared git-common-dir root.
Keys must be SHA-256 digests. Entries older than 30 days can be pruned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3SPEEGMR0DFFFXQAJJ85VWW
Both rewrites (git-refs per queued ref, git-branch over the unpushed v1 chain)
take a mode. opfScanThenApply scans anything the span cache lacks and then
applies it; the exported RewriteQueuedCheckpointRefsWithOPF and
RewriteUnpushedV1WithOPF keep that synchronous behavior. opfApplyCachedOnly
never runs the model and reports unscanned content as OPFScanPendingError,
leaving those refs untouched while covered siblings are rewritten.

The cache is resolved from the repository being rewritten, not the process's
working directory. Collected blobs now carry their object IDs.

The OPF batch cap is a pathological-input backstop again (128 MiB prose,
256 MiB raw), since the model no longer runs inside git push.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3SPSCTZEG1M9KD6357HWC8D
The pre-push hook no longer runs the OPF model. It rewrites checkpoints from
the span cache and holds back anything not scanned yet: on git-refs those
refs stay queued while scanned siblings ship, on git-branch v1 is not pushed
this time. Either way the user's push completes and a detached
`entire __opf_scan <remote>` worker is started.

The worker takes a per-repo lock for its whole run, collects the checkpoint
commits still lacking the OPF trailer on the primary backend (queued refs on
git-refs, the unpushed v1 chain on git-branch), scans what the cache lacks,
and then runs the same pre-push delivery for the remote the spawning push
named, with that push's OPF decision and non-interactive git. Units over a
cap are skipped, the loop repeats only while new work appears, and a
delivery that fails leaves everything queued for the next push.

The explicit migration push still scans synchronously, since it promises an
immediate push.

Also carries two fixes the worker depends on: SpawnDetached gives children
the null device instead of a pipe (a pipe drained by the exiting parent
killed the child on its first stderr write), and the spawn-throttle marker
moves into a shared spawnmarker package.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3SQFKMHPMBCMBFZZJFSD4G5
Pre-push printed that held refs "stay queued for the next push" while the
worker was about to push them itself, and the non-interactive and prompt text
still promised a ~30s scan during the push.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3SQMZAXBRY6KDR8D79Z5GHJ
entire status (text and --json checkpoint_opf_pending) reports how many
checkpoints are held until the background scan covers them, on both backends:
queued refs without the trailer on git-refs, unpushed untrailered v1 commits
on git-branch. The count reads local refs only and is omitted when OPF is off.

When v1 is held for a scan, git-branch users see a tip, at most once every 30
days per repository, pointing at `entire doctor migrate-checkpoints`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3SQRDXRVVA9WV1Y0VYAMXZ5
The security doc now says OPF never runs inside git push: pre-push holds
unscanned checkpoints, the scan worker scans and delivers them, the span cache
stores only hashes and offsets, and failures keep checkpoints held rather than
aborting pushes. Cap values, timeout_seconds, the prompt text and the CI note
match the new behavior.

The OPF command-trust integration tests wait for the background worker to log
that it finished before checking whether the payload ran, since that is now
where the opf binary executes; otherwise the negative tests would pass
vacuously. The worker logs that line on every exit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3SRQS5BCD884ER2E6ZX76KD

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 05e1e64. Configure here.

Comment thread cmd/entire/cli/strategy/manual_commit_push.go
Copilot stopped reviewing on behalf of peyton-alt due to an error September 30, 2026 18:45

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Copilot was unable to run its full agentic suite in this review.

Copilot review overview

Review effort: Lite
Findings: 1 High severity · 3 Medium severity · 1 Low severity

Open (5)
What changed in this PR

This PR refactors OpenAI Privacy Filter (OPF) handling to avoid running model inference on the git push hot path by introducing a persistent span cache plus a detached background scan worker, while also updating OPF batching/timeouts and exposing backlog visibility via entire status.

Changes:

  • Add an OPF span cache (git-common-dir) with scan/apply split APIs and tests.
  • Introduce a detached __opf_scan worker that scans held checkpoints and then delivers them, while pre-push rewrites from cache only.
  • Make OPF shell-out deadlines size-adaptive and chunk large batches into bounded calls; update docs and status surfaces accordingly.
File Description
redact/​opf_test.go Updates timeout-default behavior tests; adds adaptive timeout + deadline wiring tests
redact/​opf_cache_test.go Adds tests for scan-then-apply caching behavior and safety properties
redact/​opf_cache.go Introduces OPFSpanCache interface + scan/apply cached flow
redact/​opf.go Adds size-adaptive timeouts; increases raw input cap; uses adaptive timeout in shell-out
redact/​batch_test.go Adds tests for chunking logic and bounded-call splitting behavior
redact/​batch.go Refactors leaf collection; adds chunked scanning; adds NamedBlob.ID and helpers
docs/​security-and-privacy.md Updates user-facing behavior/docs for background scanning, caps, and timeouts
docs/​development/​opf-throughput-findings.md Adds benchmark findings motivating architectural changes
docs/​development/​opf-bug-investigation.md Adds investigation write-up motivating new caps/flow
cmd/​entire/​cli/​trail_context_cache.go Switches to shared spawnmarker throttling helper
cmd/​entire/​cli/​strategy/​manual_commit_push_test.go Updates per-ref delivery test to reflect cache-only pre-push and worker spawning
cmd/​entire/​cli/​strategy/​manual_commit_push.go Holds unscanned checkpoints, spawns worker, rewrites from cache-only on pre-push
cmd/​entire/​cli/​strategy/​manual_commit_opf_scan_test.go Adds tests for scan worker behavior, locking, and delivery
cmd/​entire/​cli/​strategy/​manual_commit_opf_scan.go Adds detached worker implementation + backlog counting helper
cmd/​entire/​cli/​strategy/​manual_commit_opf_rewrite_test.go Updates rewrite tests for cache-only mode and background-worker semantics
cmd/​entire/​cli/​strategy/​manual_commit_opf_rewrite.go Adds rewrite modes, pending error, cache-backed redact path, and new caps
cmd/​entire/​cli/​strategy/​manual_commit_opf_refs.go Adds mode-aware rewrite path for git-refs backend
cmd/​entire/​cli/​strategy/​manual_commit_opf_prompt_test.go Updates prompt/progress messaging expectation
cmd/​entire/​cli/​strategy/​manual_commit_opf_prompt.go Updates non-interactive progress line and prompt description
cmd/​entire/​cli/​status_test.go Adds OPF backlog visibility tests (text + JSON)
cmd/​entire/​cli/​status.go Adds OPF backlog computation and status rendering/JSON field
cmd/​entire/​cli/​spawnmarker/​spawnmarker.go Introduces shared spawn marker throttling utility
cmd/​entire/​cli/​settings/​settings.go Adds settings-level OPFEnabled helper (distinct from runtime-configured OPFEnabled)
cmd/​entire/​cli/​session_sweep.go Switches to spawnmarker throttling helper
cmd/​entire/​cli/​root.go Registers new hidden __opf_scan command
cmd/​entire/​cli/​opf_scan_cmd_test.go Tests __opf_scan wiring through root + file logging
cmd/​entire/​cli/​opf_scan_cmd.go Adds hidden __opf_scan command implementation
cmd/​entire/​cli/​integration_test/​opf_command_trust_test.go Updates integration tests to wait for background worker completion
cmd/​entire/​cli/​execx/​spawn_detached_test.go Adds test ensuring detached child stdio is null-device
cmd/​entire/​cli/​execx/​spawn_detached.go Refactors spawn into testable detachedCommand; avoids pipes that can SIGPIPE child
cmd/​entire/​cli/​checkpoint/​opf_span_cache_test.go Adds git-common-dir OPF span cache tests + pruning behavior
cmd/​entire/​cli/​checkpoint/​opf_span_cache.go Adds on-disk OPF span cache implementation + pruning

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread cmd/entire/cli/strategy/manual_commit_opf_scan.go
Comment thread redact/batch.go
Comment thread redact/opf.go
Comment thread redact/opf_cache.go
Comment thread cmd/entire/cli/execx/spawn_detached.go Outdated
peyton-alt and others added 2 commits September 30, 2026 14:35
When the OPF span cache could not be opened, pre-push reported the content as
pending, so the user saw "being scanned in the background" on every push while
the worker hit the same error and exited. It now withholds the checkpoints with
OPFCacheUnavailableError, which names the cause and the remedy, and starts no
worker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3T3WEGGV1G4VW9FCFD8MGQ6
For arbitrary blob contents, applying cached results must equal the one-pass
batch output, and an entry missing any single prose leaf must make the apply
report pending rather than redact partially.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 01M3T3WG1QND15NQ65TSXE0V20
peyton-alt and others added 13 commits October 6, 2026 10:45
…eyton/opf-scan-worker

# Conflicts:
#	cmd/entire/cli/strategy/manual_commit_opf_refs.go
#	cmd/entire/cli/strategy/manual_commit_push.go
The scan worker collected every queued ref's blobs into one slice before
scanning, so a long backlog of individually in-cap refs could grow its
memory without bound. It now walks pending commits first and loads and
scans their content in batches no larger than one ref's raw cap,
releasing each batch before collecting the next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Spawn a worker whenever none holds the lock, instead of at most once a
  minute, and have an exiting worker re-check once after releasing the
  lock, so work held by a push in either window is not stranded.
- Store scan results after each model call's worth of blobs, and deliver
  after each worker batch, so an interrupted scan keeps its progress and
  scanned checkpoints ship without waiting for the whole backlog.
- Give the OPF rewrite's remote-tip temp ref a per-process name; a
  concurrent worker delivery and user push could delete each other's
  ref, which reads as "remote has no v1".
- Skip units over the bootstrap limit in the worker, as the rewrite does,
  and clean up pushed shadow branches when v1 is held.
- Drop "(runs in the background)" from the status line, which counts
  checkpoints no worker may be scanning, and document that the cache
  version must be bumped when the model changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Entire-Checkpoint: 01M4E83KBKGT7JNM0069G3KPD0
…tself

On git-branch, holding back Entire's own v1 push does not stop the
user's push from sending the branch (`git push --all`, `--mirror`, or an
explicit refspec): git chooses what to send before the hook runs. The
pre-push line now saves git's ref list, gives it to Entire, and replays
it to whatever runs after (a chained hook or the next Husky line, e.g.
git lfs pre-push). Under OPF, a push that carries v1 at anything other
than the tip OPF just verified is refused.

The same line ends the script when Entire fails. Previously a chained
pre-push hook's exit status decided the push, so an OPF abort was lost.

The signal is ENTIRE_PRE_PUSH_STDIN_REFS=1, not a flag, so an older
binary run by the new script ignores it instead of failing every push.
Installed scripts without it read as outdated and are reinstalled; an
older hand-copied hook-manager line gets a warning when v1 is held.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Entire-Checkpoint: 01M4E9G4VHSPN4WY1JMRKQRSPN
The integration helper returned at the first "opf scan worker finished",
so a worker spawned by a later push could still be writing to .git when
the test's temp dir was removed. Log "finished" on every worker exit and
wait until each spawn has a matching finish.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The outer-push guard only covered git-branch. On git-refs, `--mirror`, an
explicit refspec, or `--all` with a v1 branch left from before a
migration could still send checkpoint content OPF had not verified. The
git-refs pre-push now refuses the user's push when it sends a checkpoint
ref (v1 or refs/entire/checkpoints/*) whose commit lacks the OPF
trailer, unless OPF is skipped for this push. On both backends a ref the
remote already has at that commit is not counted as sent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ParsePrePushRefs capped stdin at 16 MiB, so a `--mirror` of a repository
with many tags could truncate the list before the v1 line and silently
skip the outer-push check. It now reads to the end and keeps only the
lines that touch checkpoint refs, which bounds memory without a cap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Entire-Checkpoint: 01M4EBR89WEF9793069XD9F0QC
…eyton/opf-scan-worker

# Conflicts:
#	cmd/entire/cli/status.go
#	cmd/entire/cli/strategy/manual_commit_push.go
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Entire-Checkpoint: 01M4EP93BENSX3HWM79AEBQ7TG
…nned ones

The git-refs rewrite reported only its first error, so a ref over the
bootstrap or size cap hid the pending signal from an unscanned sibling
behind it. Pre-push then never started the scan worker, and the sibling
stayed held on every push while the capped ref sat first in the queue.
The rewrite now returns both, and the withheld warning prints the scan
notice and the cap error.

Also corrects two comments: the span cache is shared between the scan
worker and pre-push, not used by one process at a time, and the detached
child example names `__opf_scan` rather than the removed `__opf_flush`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Entire-Checkpoint: 01M4F6C6RP64X6BW8SF0QKVRSY
The previous fix joined the pending signal onto a cap error queued ahead
of it, but a pending ref queued first still became the only error and
the cap was dropped, so the user saw "scanning in the background" for a
ref that will never ship. The rewrite now keeps pending apart from every
other error and returns both, in either order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Entire-Checkpoint: 01M4F79SK3NGWYGRER6RXM7SVX
…rker

- The pre-push line falls back to running Entire without the ref list
  when the temp file cannot be created or written, and an unreadable
  list is treated as unknown instead of failing the push.
- A scan worker stops once the settings say OPF should never run.
- `entire clean --all` removes the OPF span cache.
- The span cache key includes the configured OPF command.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Entire-Checkpoint: 01M4GY0FJRX06GEEPFAA7YX2QN
…ng it

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants