Skip to content

Latest commit

 

History

History
323 lines (260 loc) · 15.6 KB

File metadata and controls

323 lines (260 loc) · 15.6 KB

Extending codeindex-sync

Two extension points, in increasing order of effort:

  1. Add a backend — usually pure config, no code.
  2. Add a hook handler — react to git activity for something other than indexing.

1. Add an index backend

If your backend is an MCP server exposing index/update/status tools, you do not write code. Describe it in ~/.config/codeindex-sync/config.json:

{
  "providers": [
    {
      "name": "my-backend",
      "description": "What it does, one line",
      "command": "my-index-server",
      "args": ["--stdio"],
      "tools": {
        "update": "update_index",
        "index": "rebuild_index",
        "status": "index_status"
      },
      "repoArg": "projectPath",
      "detectFiles": [".my-backend.json"],
      "busyMarkers": ["already indexing"],
      "timeoutMs": 3600000,
      "env": { "MY_BACKEND_URL": "https://..." }
    }
  ]
}

codeindex-sync providers --example prints a starting block.

The fields that matter

tools.update is the only required tool. Everything else degrades: without index, a --full request falls back to update; without status, list reports nothing for that provider rather than failing — and an asynchronous index cannot be verified, so configure status if you configure index.

tools.list is what makes tools.remove trustworthy. A removal is confirmed against the backend's own listing, so configuring remove without list leaves cleanup --apply able to ask but not to check — it will say so rather than claim a removal it cannot see. See below.

detectFiles decides routing. When several providers are configured, the first one claiming a repo wins, in config order. A marker file is the usual mechanism — it also lets a repo opt in explicitly. An empty list means "claim any repo", which is fine when only one provider is configured.

markerContent is what lets claim write that marker for you. The file name is already detectFiles[0]; this is what goes inside it, and that part is irreducibly backend-specific — SocratiCode pins a collection id, yours may want something else. ${name} expands to the repository's directory name. Leave it out and claim says it cannot write the marker rather than inventing a format your backend will not understand.

busyMarkers prevents a real bug. Most backends hold their own per-project lock. When another indexer holds it, the reply is contention, not failure — the job must be requeued without burning a retry attempt. Get this wrong and three unlucky collisions park a perfectly healthy repository in failed/. Markers are matched as whole words or phrases, case-insensitively, only against what the backend actually replied and with the repository's own path removed first — not against the error text of a session that died. Prefer phrases a backend says ("already indexing this project") to bare words that can turn up in a file name.

Retries are top-level config, not per provider. A failed attempt is retried in the same process — by sync, once and drain alike — after backoffSeconds (default 10), doubling each time, until maxAttempts (default 3, the first included) have been made. The job is then parked in failed/, where status shows why and retry requeues it.

timeoutMs and killAfterMs bound a backend's life. timeoutMs (default one hour) caps a whole session; a first full index of a large repository needs more. However a session ends, the backend's input is closed, it is sent SIGTERM once, and then it is waited for: a backend that finishes its in-flight work before exiting still holds its project lock until it does, and a retry started beside it would be refused that lock. killAfterMs (default 75 seconds) is how long the wait may last before the backend is killed outright. Each backend runs in its own process group: SIGTERM goes to the launcher first — npx passes it on, and a backend should not get it twice — then to whatever is left of the group once the launcher has exited, and the kill goes to the whole group, so a backend behind a launcher or a shell script is reached either way. For the same reason the CLI passes a SIGINT, SIGTERM or SIGHUP that interrupts it on to any backend still running; code embedding McpSession can opt into the same with forwardSignalsToBackends(). A run whose backend had to be killed is parked at once rather than retried, because its lock outlives it until the backend's own staleness rule frees it. A run that merely timed out is retried like any other failure: once its backend has exited, a backend that checkpoints resumes where the last attempt stopped.

asyncIndexMarkers prevents a worse one. A full-index tool is often fire-and-forget: it starts the work on the backend's own event loop and returns in about a second saying so. Taken at face value that reply is a silent data-loss bug — the session closes, the child is reaped moments into a job needing minutes, and an empty index reports as a success.

So a reply is not evidence the work happened. After invoking tools.index, codeindex-sync polls tools.status on the same session until the backend stops reporting progress, and reports what status says. One session, because progress is usually per-process state: a second child would see a backend that has never indexed anything.

Three fields tune it, and the defaults suit any backend that says the usual things ("in the background", "in progress"):

Field Default What it matches
asyncIndexMarkers in the background, running asynchronously, check progress A reply meaning "started", not "done"
progressMarkers in progress, in-progress, actively indexing A status reply meaning work is still running
pollIntervalMs 2000 Cadence the status polls settle at (must be positive)

Polling starts immediately and backs off to pollIntervalMs, so a tool that already did the work before replying pays one extra call and no waiting, while a long job is not polled hard. Cheap early polls also matter for correctness: a run that starts and finishes between two polls was never seen, and an unseen run cannot be told apart from one that never started.

Three consequences worth knowing:

  • An incremental update that answers synchronously is not polled — hooks fire constantly and that reply is already the truth — but one whose reply matches asyncIndexMarkers is.
  • "Done" is not merely "no progress marker". Re-indexing a repo that already has an index reports a perfectly healthy one during the window before the new run becomes visible, so a reply that announced background work must be watched running before it counts as finished.
  • A run that stops having produced no index at all, or one the backend still calls incomplete, is a failure — not a success with a small number in it.

A status tool that answers with an error is not a verdict. Its backend is running, and the error can be the backend's own read tripping over storage it is creating at that moment — SocratiCode's codebase_status fails exactly this way when a poll lands while a first index is still creating its Qdrant collection. Abandoning the run over it would close the session and kill the index being waited on. So the status tool is asked again, and only errors that go on unbroken for the settle window (eight pollIntervalMs, 16 seconds by default) end the wait as a failure. A session that died, timed out, or rejected the call outright — an unknown tool, bad arguments — ends it at once, because nothing more can be learned from it.

Beyond that, the wait is bounded by the session's hard timer, which ends the session and makes every later call fail at once, and — for code that embeds the provider — by an abort signal. Both end as a failure: a backend that never finishes must never look like one that did.

A removal is verified, not assumed. cleanup's entire input is indexes whose directory is gone — that is the definition of an orphan — and a backend that identifies an index by a marker inside the repository cannot resolve it once the directory is deleted. It then removes nothing, has nothing to complain about, and answers without an error. Believing that reply is how cleanup --apply printed ✔ removed while the index stayed put and re-appeared as an orphan on the very next run, forever.

So remove() calls the tool and then re-reads tools.list, reporting one of three things — and only the first is shown as removed:

Outcome Meaning
removed The backend no longer lists the index
failed The tool errored, or the index is still listed afterwards
unverified The tool was accepted, but no tools.list was available to confirm it

Unlike a slow index call, this check does not need to share a session: a project listing is durable backend state rather than per-process progress, so asking again in a fresh process is stricter, not weaker — it proves the removal is visible to the next process, which is exactly what the next cleanup run will be.

When a removal fails this way, the fix is usually to give the backend back the marker it needs: recreate the directory with just that file (detectFiles names it) and remove the index with the backend's own tooling. cleanup prints that advice, built from your detectFiles.

env exists because git hooks are not a login shell. They never source ~/.bashrc, ~/.zshenv or any profile, so anything you export in a shell is invisible to the indexer. If your backend needs configuration, it goes here or in a file the hook path reads directly — never in a shell profile.

Your backend's log lands in the worker log. A server that declares the MCP logging capability sends its log lines as notifications/message, and every one it sends during an index run is written to the log, and to the output of sync, drain and once, as [<name>:<level>] <message>. MCP's eight levels become four: debug; info (and notice); warn; error (and critical, alert, emergency). A level it does not recognise is logged as info. Data that is not a string is written as JSON, line breaks become |, and a line is capped at 2,000 characters, with a note of how much was cut. The backend's logger name is added in front of the message when it differs from the provider's name.

Nothing is filtered here: codeindex-sync has no verbosity setting of its own, so the backend's configured log level decides what arrives, and no logging/setLevel is sent to override it. Lines can arrive at any point in a session, including before the backend answers initialize. Only index runs are logged, so list, doctor and cleanup drop them as before.

One thing config cannot describe: integrity

codeindex-sync verify is the single command that does not go through the provider interface. It compares what an index claims to hold — the file-hash map the backend keeps — against what it actually holds, and no MCP tool exposes either side, so it reads the backend's store directly.

That makes it backend-specific by construction, and it is deliberately fenced off: src/qdrant.ts and src/verify.ts are the only files that know a storage engine exists, and nothing in the worker, the queue or the provider interface imports them. Every other command still works against any MCP backend.

Supporting a different store would mean a second implementation behind the same two questions — "what does this index claim?" and "what does it hold?" — rather than a change to IndexProvider. A backend that grows a tool answering those would be better still, and would make this command config-only like the rest.

If config isn't enough

A backend that needs real code is a signal the IndexProvider interface is missing something. Prefer extending the interface over special-casing a product, so the abstraction keeps its guarantee: nothing above provider.ts knows any backend's name.

To implement one anyway:

import type { IndexProvider } from "codeindex-sync";

export class MyProvider implements IndexProvider {
  readonly name = "mine";
  readonly description = "…";

  async detect(repoPath: string) { /* cheap, no network */ return true; }
  async index(req) { /* → { status: "ok" | "busy" | "failed", summary } */ }
  async health() { /* → ProviderHealth[]; must never throw */ return []; }
  async status(repoPath: string) { return null; }
  // Optional. Must confirm against the backend before reporting "removed".
  // → { status: "removed" } | { status: "unverified" | "failed", detail: string }
  async remove(repoPath: string) { /* must confirm before reporting "removed" */ }
}

Three rules, each learned from a production failure:

  • detect must be cheap. It runs for every provider on every job.
  • health must never throw. It backs doctor, which has to work when everything else is broken. Return ok: false with a remedy instead.
  • Distinguish busy from failed. See above.

If your backend keeps a log of its own, pass it on through req.log, one line at a time: req.log?.({ level: "warn", message: "…" }), where level is debug, info, warn or error. The worker writes it as [<name>:<level>] <message>. Keep each message to one line; the log is read line by line.


2. Add a git-hook handler

Indexing is only the first subscriber to git activity. A handler can do anything — warm a lint cache, regenerate docs, notify something:

import { HookRegistry, type HookHandler } from "codeindex-sync";

const lintCache: HookHandler = {
  name: "lint-cache",
  description: "Warm the lint cache after a checkout",
  hooks: ["post-checkout", "post-merge"],
  async handle(event) {
    // event.repoPath is the project root, already resolved
    // event.hook, event.args, event.at
  },
};

registry.register(lintCache);

Rules for handlers

Be fast, or defer. Hooks run inside the user's git command. One git rebase can fire dozens of events. Enqueue work; don't do it inline.

Never rely on the environment. Hooks are not a login shell.

Use event.repoPath, never process.cwd(). The hook's working directory is frequently a throwaway worktree that no longer exists by the time you run — and a deleted cwd kills the process during interpreter startup, before your code runs at all. repoPath is already resolved: the nearest enclosing directory a configured provider claims, else the repository root.

Some hooks never reach you, deliberately. A hook firing in a linked worktree, or in a directory matching excludePaths (agent tools: .claude/worktrees/, .codex/worktrees/), is dropped before dispatch — no handler is called. Those commits change no file in the checkout that carries the index, and the directory is usually gone seconds later. If your handler genuinely wants worktree events, say so on an issue; nothing about the design forbids it, but the guard is currently before the registry rather than inside it.

Throwing is contained but not free. A handler that throws is reported and the others still run — a broken extension must never break git commit. But a slow one delays every git command the user types.

Per-handler timings come back from dispatch(), so a slow extension is identifiable rather than just "git feels sluggish lately".


Testing your extension

Both interfaces are plain objects, so no framework is needed:

const provider = new MyProvider();
expect(await provider.detect("/repo")).toBe(true);

The worker takes an injected provider, so end-to-end behaviour — retries, coalescing, crash recovery — can be tested against a scripted fake with no real backend running. See test/worker.test.ts for the pattern.