A Claude Code plugin that turns a Zotero library and your agent session logs into a maintained wiki — one where every assertion traces to a source, and the machinery that could quietly break that chain is fenced off deterministically rather than by instruction.
A reference manager knows what you have. A note pile knows what you wrote. Neither can tell you where your sources disagree, which concept rests on a single paper, or what you own and have not read that bears on the question you are stuck on.
Datalinks builds that layer — and, more importantly, refuses to build it out of things it cannot verify.
The failure it is designed against is specific: an LLM wiki that cites its own past reasoning as literature. After a few cycles it becomes a confident, densely cross-referenced record of things nobody checked — and every citation resolves, so every integrity check passes. That failure is invisible once it has happened. Most of what follows exists to make it impossible.
Alpha. Complete and self-consistent — 11 skills, 5 agents, 7 scripts — with CI covering component conformance and an end-to-end fixture run. Not yet exercised against a large real corpus. Interfaces may move.
git clone https://github.com/exarcos/datalinks
cd datalinks
pip install pyyaml
uv tool install zotero-mcp-server # or: pipx install zotero-mcp-serverIn Zotero: Settings → Advanced → “Allow other applications on this computer to communicate with Zotero.” Better BibTeX is recommended — it supplies the stable citekeys used as display names, and CSL-JSON for export.
export DATALINKS_VAULT=~/vaults/research
export DATALINKS_DB=~/.local/share/datalinks/index.db # NOT inside the vault
export DATALINKS_FULLTEXT="$DATALINKS_VAULT/sources/.fulltext"
export DATALINKS_REDACTION_SALT=$(openssl rand -hex 16)
mkdir -p "$DATALINKS_VAULT"/{sources,sessions,notes,_datalinks/{pins,rubrics}}
mkdir -p "$DATALINKS_VAULT"/wiki/{claims,concepts,topics,definitions,questions,deductions}
cp rubrics/*.yaml "$DATALINKS_VAULT/_datalinks/rubrics/"
git -C "$DATALINKS_VAULT" initImportant
Keep the index database outside any synced folder. SQLite inside iCloud, Dropbox or Obsidian Sync will eventually be corrupted by concurrent partial-file sync — Zotero warns about the same hazard for its own database.
Then, in Claude Code:
pull in the Dissertation/Ch3 collection from Zotero
extract claims from the new annotations
regenerate the concept pages
run the gardener
flowchart LR
Z[("Zotero<br/>library")] -->|wiki-ingest| S["sources/<br/><i>pure projection</i>"]
L[("Agent<br/>session logs")] -->|redact.py| SE["sessions/"]
S --> X{{"rubric<br/>gate → extract"}}
SE --> X
X --> C["claims · definitions<br/>open questions"]
X --> D["decisions · procedures"]
C -->|wiki-synthesize| W["concepts · topics"]
C --> G["wiki-garden<br/><i>adjudicate · relink<br/>close · deduce</i>"]
G --> C
W -->|wiki-export| O["document · bibliography<br/>glossary · site"]
N["notes/<br/><i>yours alone</i>"] -.cited by.-> C
P["pins<br/><i>your corrections</i>"] -.constrain.-> W
Sources are projections. No model output ever enters a source note, so
deleting sources/ and regenerating produces byte-identical files. Everything
downstream inherits that stability.
Extraction is rubric-driven. Nothing enters because a model found it interesting — it enters because a rubric you own routed it there, with every criterion verdict recorded, including rejections.
Judgment is spent only where judgment is needed. Quote verification is a string search first, so most citations settle for free and only near-misses reach a model. Secret redaction never uses a model at all — partly because a model misses things unreproducibly, and mostly because routing a secret through a model to ask whether it is a secret has already leaked it.
Maintenance is bounded. The gardener runs unattended with per-phase change budgets that abort rather than truncate, deletes nothing, and must write nothing on a second consecutive run.
| Finding | Answers |
|---|---|
unmined_neighborhood |
You own 12 unread items bearing on your most contested concept |
contested_hotspot |
This is where your sources actually disagree |
single_source_concept |
All 7 claims here rest on one paper |
research_gap |
Many questions, few sources — under-read. Or: many questions despite many sources — unsettled |
collection_coverage |
Which of your own buckets have you actually mined? |
| Ask for | Skill |
|---|---|
| "pull in this collection from Zotero" | wiki-ingest |
| "extract claims from these annotations" | wiki-extract |
| "regenerate the concept pages" | wiki-synthesize |
| "run the gardener" / scheduled maintenance | wiki-garden |
| "audit the wiki" / before an export that matters | wiki-audit |
| "export this topic" / "make me a bibliography" | wiki-export |
| "record this decision" | wiki-record |
| "review my tags" | wiki-taxonomist |
| "which collections have I actually read" | wiki-collections |
| "switch reference managers" | wiki-rebind |
| anything about schema, evidence rules, IDs | datalinks-conventions |
Scripts run standalone, no Claude required:
python3 scripts/reindex.py --full # rebuild the index
python3 scripts/reindex.py --verify # detect a drifted index
python3 scripts/validate.py --corpus # schema, evidence matrix, provenance
python3 scripts/verify_quotes.py # verbatim citation checking
python3 scripts/export.py --scope topic:Networks --target document --dry-run
python3 scripts/redact.py transcript.jsonl --jsonl --report
python3 scripts/migrate.py --statusSix invariants, stated once in
datalinks-conventions and enforced by
validate.py. Three are worth knowing before you write
anything:
A session is never evidence that a claim is true. It is evidence of what was decided. Rejected at write time, not warned about.
Ownership is by path. The agent writes everywhere except notes/. Your
corrections to generated pages become pins — structured judgments re-applied
on every regeneration — rather than edits the next rebuild destroys.
Propose, never apply, outside the generated layer. An unapplied proposal is visible and costs a review; a wrongly applied change is invisible and costs trust.
None required; each is declared as a capability, so dependent features become unavailable rather than broken when absent.
| Provider | Adds |
|---|---|
| Academix | Discovery and citation-graph edges across OpenAlex, Crossref, Semantic Scholar, arXiv, DBLP |
Crossref, or zotero-mcp's Scite extra |
Retraction status |
| mcp-refchecker | Verifies a proposed source exists |
| marker / docling | Page-segmented full text |
Wire up retraction status; leave the rest until the wiki is developed enough for discovery to pay off.
Note
Known gap. verify_quotes.py confirms a quote appears in its source but not
that it appears on the cited page — Zotero's full-text index has no page
boundaries. A correct quote with a wrong locator passes every check here.
Closing it needs page-segmented extraction.
| Variable | Purpose |
|---|---|
DATALINKS_VAULT |
Vault root. Unset disables every hook silently. |
DATALINKS_DB |
Index database. Keep it out of synced folders. |
DATALINKS_FULLTEXT |
Cached source text for quote verification. |
DATALINKS_REDACTION_SALT |
Makes redaction markers stable across sessions. |
DATALINKS_ALLOW_DELETE |
Permits deletion inside the vault. Deliberate regeneration only. |
The full scoping document — phased build plan, the port/adapter split for
alternative reference managers and vaults, the corpus-level deduction catalogue,
and the reasoning behind each invariant — is in
docs/design.md.
Python 3.11+, PyYAML. Optional: pandoc for document export, Zotero 7 with the local API enabled.
Datalinks ships as a Claude Code plugin, but its core is harness-independent — the
scripts are plain Python and the Zotero source is reached over MCP, which Claude
Code, OpenAI Codex, ChatGPT developer mode, Gemini CLI and Cursor all speak. A
root AGENTS.md is the portable standing brief those harnesses read,
and MCP config ships in each format they require (.codex/config.toml,
.gemini/settings.json, .mcp.json).
The one thing that does not port is hook-based enforcement: Claude Code prevents
invariant violations at the tool boundary, while other harnesses must uphold the
same rules by running validate.py and verify_quotes.py in-loop and treating
their exit codes as gates. docs/harnesses.md covers which
harness consumes which file and where that safety margin narrows.
See CONTRIBUTING.md. Run scripts/selfcheck.py and
tests/smoke.sh before opening a PR — CI runs exactly those.
The wiki-as-compounding-artifact framing follows
Karpathy's LLM Wiki;
§12 of the design document records where this implementation agrees with that
spec and where it deliberately does not. Zotero integration builds on
zotero-mcp.