Skip to content

cli(cache): a metadata-only file change rejects the unchanged candidate — every later command for that project fails until the cache is cleared #4

Description

@EnRaiha

Version / build tested against

Identical result on both. code2graph 0.0.0, debug profile.

Deployment mode

CLI, local cache. Cache file at $XDG_CACHE_HOME/code2graph/projects/<project-key>/cache.sqlite3. Single process, sequential runs, no concurrency.

Engine(s) involved

code2graph-cli cache publication (cli/src/cache). Extraction, resolution and query paths are unaffected.

Summary

A file whose content is unchanged but whose mtime moved cannot be published. The candidate id addresses content, while the stored candidate row carries mtime, so publication finds its own earlier candidate, compares metadata, and rejects the write. From then on every command for that project fails, including read-only queries, until the cache directory is cleared.

The error names an internal invariant and no recovery.

Downstream damage:

effect detail
trigger an ordinary touch, cp, rsync, checkout, restore, or build step
blast radius index, symbols, def, status, module-deps. The project cache is unusable, not merely stale
self-healing none. The transaction that would refresh the stored metadata is the one the conflict aborts

Steps to reproduce

mkdir -p /tmp/repro && cd /tmp/repro
printf 'fn a() {}\n' > a.rs
printf 'fn b() {}\n' > b.rs
printf 'fn c() {}\n' > c.rs

c2g index --root /tmp/repro          # ok
touch /tmp/repro/*.rs                # content byte-identical, only mtime moves
c2g index --root /tmp/repro          # fails

Observed:

{"schemaVersion":1,"status":"error","error":"cache failure: cache candidate conflicts with an existing candidate id"}

Every later command on that project reports the same error:

c2g symbols a --root /tmp/repro      # cache failure: cache candidate conflicts with an existing candidate id
c2g status --root /tmp/repro         # cache failure: cache candidate conflicts with an existing candidate id

Recovery, and the only one:

c2g cache clear --root /tmp/repro    # {"status":"ok","operation":"clear","scope":"project","removed_projects":1,...}
c2g index --root /tmp/repro          # ok again

Expected behavior

A metadata-only change either mints a new candidate or refreshes the stored metadata in place. Re-running any command on the same tree succeeds, and the cache keeps serving the project.

Actual behavior

Publication fails, and the project cache is left in a state no later run can leave. The asymmetry is visible in one command: c2g index --root /tmp/repro --no-cache succeeds on the same tree, because it never publishes.

What actually happened? (check all that are true)

  • data loss / corruption
  • crash / hang
  • security
  • core broken: the project cache, and every command that reads it
  • workaround exists: c2g cache clear --root <dir> discards the project cache and forces a full re-index, or --no-cache avoids publication for a single run

Proposed severity

SEV-3 — Medium: feature wrong, but operational and a workaround exists.

Proposed labels: type:bug, sev:3-medium, status:needs-triage. Not regression; see below.

A maintainer may confirm sev:2-high instead: the failure is total for the project, the trigger is an ordinary metadata change in the edit-compile loop the cache exists for, and the error names no recovery.

Reproducibility

Always. Three files, one touch, fresh cache directory. Reproduced on four fixtures across the two commits above, and once more on a 300-file project.

Last known-good version / commit (if a regression)

Unknown. The identity and verifier pair dates from 0975983 ("Adopt metadata-first refresh planning", 2026-07-11). I did not test a build predating it, so regression is not claimed.

Environment & logs

Debian 13 (trixie), Linux x86_64, cargo build debug profile, code2graph 0.0.0. Cache under a scratch $XDG_CACHE_HOME. All output above is quoted verbatim.

Before submitting


Root cause

Read from source, not inferred:

  • The input digest hashes (path, language, content_hash) only, so identical bytes produce an identical digest across a touch. cli/src/cache/fingerprint.rs:172
  • CandidateId::new(compatibility, input_digest, completeness, &omissions) is therefore content-addressed, not metadata-addressed.
  • The persisted file row carries mtime: INSERT INTO candidate_files (... mtime_seconds, mtime_nanoseconds ...). cli/src/cache/store.rs:423
  • On publication, verify_existing_candidate rejects a differing mtime. cli/src/cache/store.rs:1238

Same content plus a new mtime yields the same candidate id with a different payload, so the fourth step fires on every later run. The failure returns Err from the BEGIN IMMEDIATE body and the transaction rolls back (cli/src/cache/store.rs:587), so the stored mtime is never refreshed. The dead end is permanent by construction.

Identity and verifier disagree about whether mtime identifies a candidate. The refresh planner expects the second answer. By default a file with a prior record is planned as NeedHash, and --trust-mtime reuses facts only when size, mtime, language, tier and package all match, falling back to a hash otherwise (cli/src/refresh/plan.rs:142, :169). A metadata-only change is designed to fall back to a content hash, not to fail. The planner takes the change; the verifier refuses to publish the result.

The trigger boundary confirms it. Only the metadata moves:

scenario content mtime result
edit the file changed moved ok
touch identical forward fails
touch -d 2020-01-01 identical backward fails
touch, then --trust-mtime identical forward fails
nothing changed identical identical ok

Fix directions

  1. Put mtime in the identity. A touch then mints a new candidate. Correct, and one candidate row per metadata-only change, with the graph republished.
  2. Keep mtime out of the identity and treat it as refresh metadata. Tolerate a differing mtime in verify_existing_candidate and refresh the stored value from the incoming candidate. The same function already fills a missing file_subgraph with an UPDATE rather than rejecting it.
  3. Whatever the choice, make the failure actionable. A conflict between identity and payload must not be a dead end. Naming the file and the field that differ would make this diagnosable from outside; today the operator gets an internal invariant and no next action.

Happy to test a patch, or to run the reproduction against a large repository (--trust-mtime and multi-tier runs included) to confirm the fix holds.

Activity

  1. changed the title [-]A metadata-only file change (touch) wedges the project cache: every later command fails with "cache candidate conflicts with an existing candidate id"[/-] [+]cli(cache): a metadata-only file change rejects the unchanged candidate — every later command for that project fails until the cache is cleared[/+] on Oct 4, 2026
  2. farhan-syah commented on Oct 4, 2026

    @farhan-syah
    Member

    Okay, i'll check it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions