A reproducible multi-agent reviewer for academic economics papers, using Codex or Claude Code. The repository contains the preprocessing, reviewer prompts, validation, normalization, editor assembly, and tests. Papers and generated reports remain local and are not included.
For each paper, the wrapper:
- preprocesses the PDF locally into source-faithful text, coordinates, page images, tables, figures, citations, and cross-references
- runs a parser-quality preflight before substantive review
- selects every reviewer whose remit is plausibly relevant, with a full-roster fallback when applicability is uncertain
- runs the reviewer panel and validates every structured JSON result
- conservatively normalizes clear duplicates without discarding source evidence
- asks the editor to lead with concrete correctness and auditability problems, followed by broader positioning and development suggestions
- writes and smoke-checks the final report at
outputs/<paper_id>/report.md
The PDF parser is deterministic and local. It does not use an external document service, hosted OCR, or an LLM repair layer, and it never invents missing signs, values, labels, cells, or formulas. Unsafe artifacts are flagged so reviewers can use page images or return cannot_verify.
The review itself is not fully local: parsed manuscript text is sent to OpenAI when using Codex, or to Anthropic when using Claude Code. Search-enabled reviewers may also send manuscript-derived queries to web search. Do not review a confidential paper unless its disclosure terms permit those transmissions.
git clone https://github.com/Ingar30/reviewer.git
cd reviewerGit is convenient but not required. You can also download the repository as a ZIP and open a shell in the extracted folder.
You need:
- Python 3.12 or newer
- One authenticated CLI: Codex or Claude Code (you do not need both)
- Access to your chosen model and the CLI's web-search tools
| Backend | Sign in | Default model | Lower-cost option |
|---|---|---|---|
| Codex | codex login |
GPT-6.1 Sol | GPT-6 Luna |
| Claude Code | claude auth login |
Opus 5.5 | Sonnet 5.5 |
Use an up-to-date CLI. See model options and Claude setup for version requirements and lower-cost choices.
Windows PowerShell:
.\setup.ps1
.\.venv\Scripts\Activate.ps1macOS/Linux:
bash setup.sh
source .venv/bin/activateWindows PowerShell:
.\.venv\Scripts\python.exe -m unittest
.\.venv\Scripts\python.exe scripts\check_environment.pymacOS/Linux:
./.venv/bin/python -m unittest
./.venv/bin/python scripts/check_environment.pyFor Claude, add --backend claude to the environment-check command. These checks
do not save a preference; choose your backend when starting your first review.
Put the source PDF in inputs/. The directory is ignored by Git.
inputs/my-paper.pdf
Choose one command, using the Python environment activated above.
Codex:
python scripts/review_paper.py --backend codex --pdf "inputs/my-paper.pdf"Claude Code:
python scripts/review_paper.py --backend claude --pdf "inputs/my-paper.pdf"The backend choice is remembered for new papers in this workspace (the repository
folder for this workflow). You can omit --backend next time, or choose explicitly
for another new paper. Existing papers keep their original backend when resumed.
Commands with no saved choice still use Codex, so existing scripts keep working.
The selected backend and model are printed at startup; Reviewer never switches
providers automatically. --help and prerequisite checks do not change your preference.
The final report is written to:
outputs/my-paper/report.md
Intermediate parsed artifacts, prompts, logs, reviewer outputs, routing decisions, and the editor bundle are written to:
work/my-paper/
Use an explicit paper ID when needed:
python scripts/review_paper.py --pdf "inputs/my-paper.pdf" --paper-id "my-custom-id"The Quick Start above remains the primary workflow; uv is optional. uvx
comes with uv and sets up Reviewer and its Python dependencies in an isolated
environment, so you do not need to clone this repository or create a virtual
environment yourself.
One-time setup: you still need Git, Python 3.12+, and your chosen CLI installed and authenticated (see prerequisites). Install uv on Windows using PowerShell:
winget install --id astral-sh.uv -eFor macOS/Linux or other installation methods, see the
uv installation guide.
Open a new terminal after installation and run uvx --version to check that it
is available.
Start a review: open a terminal in the folder containing paper.pdf and
choose one command below. Replace "paper.pdf" with your PDF's filename or its
full quoted path; the folder does not need to contain the Reviewer repository.
Codex:
uvx --from git+https://github.com/Ingar30/reviewer.git economics-paper-reviewer --backend codex --pdf "paper.pdf"
Claude Code:
uvx --from git+https://github.com/Ingar30/reviewer.git economics-paper-reviewer --backend claude --pdf "paper.pdf"
For lower-cost models, see model options or Claude model choices.
This uses the same review pipeline and your chosen CLI account, consumes normal review
quota, and needs no new API key. Replace --pdf "paper.pdf" with --help or
--check for checks without a review; --check does not verify model access or
remaining quota. The manuscript-transmission notice above still applies.
Input paths resolve from your current directory, which need not be a Git repo.
Papers, intermediates and reports persist in ./reviewer-workspace/, outside uv's
cache; choose another location with --workspace "D:/Review work". The report is
at outputs/<paper_id>/report.md inside that workspace. Back it up; use distinct
paper IDs or workspaces for papers with the same filename stem.
The backend choice is remembered in that workspace, so later commands can omit
--backend. The original command without this flag still works and uses Codex
when no choice has been saved.
Existing review flags still work. Resume with the same workspace, paper ID and
runtime: new runs support --resume-incomplete to reuse validated completed
reviewers after preflight/routing; --resume-after-preflight reuses preflight but reruns substantive reviews;
--refresh-editor --paper-id paper --run-editor reuses completed reviewer outputs.
Keep the original wheel or pin a Git commit (...reviewer.git@COMMIT) for recovery.
Changed runtime resources and unrelated nonempty workspaces are not overwritten.
Codex sandbox, approval and trust rules are unchanged; no bypass is enabled.
Tested on Windows x64/Python 3.12, including live reviews. A sleep-interrupted GPT-6 run required manual recovery; unattended recovery is not established. Other platforms have not had live-review validation here. See validation and local-package testing.
--backend claude uses the same pipeline, review prompts and validators,
through Claude Code subscription login (no API key).
It defaults to Opus 5.5 (claude-opus-5-5); install Claude Code 2.1.280+ and sign in
with claude auth login. Check prerequisites without a review:
python scripts/check_environment.py --backend claudeAdd --backend claude to the ordinary review command, or use the optional uv launcher:
uvx --from git+https://github.com/Ingar30/reviewer.git economics-paper-reviewer --backend claude --pdf "paper.pdf"
With uv, results persist in ./reviewer-workspace/; --workspace DIR is optional.
Use a new paper ID or workspace for backend comparisons. Resumes reuse the saved backend.
For a lower-cost Claude option, choose Sonnet 5.5 with --model:
python scripts/review_paper.py --backend claude --model claude-sonnet-5-5 --pdf "inputs/my-paper.pdf"Or with the optional uv launcher:
uvx --from git+https://github.com/Ingar30/reviewer.git economics-paper-reviewer --backend claude --model claude-sonnet-5-5 --pdf "paper.pdf"
Sonnet 5.5 needs Claude Code 2.1.284+; run claude update if needed.
It generally uses less subscription allowance than Opus, but its review quality has not yet
been benchmarked here. Opus remains the default.
Reviews send your paper to Claude and use your subscription allowance. Keep
Usage credits / Extra usage off to stay within your subscription.
See model choices, setup and resuming a review.
Each backend has its own default: GPT-6.1 Sol (gpt-6.1-sol) for Codex and Opus 5.5 (claude-opus-5-5) for Claude. Both use xhigh reasoning for substantive reviewers and the editor and high for parser-quality preflight and applicability routing. The model notes describe the available comparisons. The GPT-6.1 Sol live test used Codex CLI 0.159.0; update an older CLI if the model is unavailable.
For a lower-cost Codex option, explicitly select GPT-6 Luna (gpt-6-luna):
python scripts/review_paper.py --backend codex --pdf "inputs/my-paper.pdf" --model gpt-6-luna --reasoning-effort xhighLuna is the lowest-cost tested Codex option, but not an equivalent-coverage replacement.
In a small three-reviewer comparison, Luna at xhigh used about 19 times fewer
estimated Standard credits than GPT-6.1 Sol at high, but missed more source-checked
issues. For broader coverage with less reasoning than the default, try Sol with
--reasoning-effort high. See model comparisons
for the limited evidence; this was not a full-pipeline comparison.
To deliberately use another combination:
python scripts/review_paper.py --pdf "inputs/my-paper.pdf" --model MODEL_ID --reasoning-effort EFFORTSupported effort values are none, low, medium, high, xhigh, and max. Overrides apply only to that run and should be treated as unbenchmarked unless evaluated on representative papers. The run manifest records the effective model and reasoning settings.
The wrapper runs up to four reviewer agents concurrently and records the PDF hash, effective model, reasoning settings, active roster, Git state, and elapsed time in work/<paper_id>/run_manifest.json.
For a run started with this version that completed preflight and routing, reuse validated completed reviewers and retry only unfinished stages:
python scripts/review_paper.py --pdf "inputs/my-paper.pdf" --paper-id "my-paper" --resume-incompleteKeep the original runtime, PDF, model, effort settings and workspace. Changed inputs or accepted outputs are refused; this does not upgrade older checkpoints. The backend is recovered from the saved run. There is no automatic quota retry or paid fallback. If preflight itself did not finish, repeat the original command.
If a run stops after a valid parser-quality preflight, resume without rerunning that stage:
python scripts/review_paper.py --pdf "inputs/my-paper.pdf" --paper-id "my-paper" --resume-after-preflightIf all selected reviewer JSON files already exist, rebuild the editor inputs and rerun only the editor:
python scripts/refresh_editor.py --paper-id "my-paper" --run-editorBoth helpers validate the saved PDF hash and existing artifacts before reuse.
Parser-quality preflight runs first. Eight universal reviewers then cover:
- cross-references
- source consistency
- claim-evidence alignment
- literature and positioning
- reference integrity
- grammar and copyediting
- abstract/conclusion consistency
- models, equations, notation, and estimands
Eleven conditional specialists cover numerical checks, identification, robustness, sample construction, limitations and external validity, theory logic, data availability and replication, institutional context, power and multiple testing, design and randomization, and economic magnitude.
A specialist is skipped only when a high-confidence classification shows that its entire remit is absent. Mixed, unknown, or lower-confidence papers use the full 19-reviewer substantive roster. Literature and reference verification use web search and return cannot_verify when evidence is unavailable.
Reviewers are configured in config/reviewers.json.
Tracked project machinery:
.codex/config.toml: project model and reasoning defaults.agents/skills/paper-reviewer/SKILL.md: reusable workflow playbookconfig/reviewers.json: reviewer roster and metadataprompts/templates/: reviewer, routing, and editor promptsschemas/: structured output contractsscripts/: preprocessing, orchestration, validation, normalization, and report checkstests/: focused regression testsdocs/first_review_walkthrough.md: step-by-step first-run guidedocs/extension_guide.md: extension points for forks
Local/private runtime locations:
inputs/: source PDFswork/<paper_id>/parsed/: deterministic parsed artifactswork/<paper_id>/selection/: applicability decision and active rosterwork/<paper_id>/reviews/: reviewer JSONwork/<paper_id>/editor/: normalized bundle and editor inputoutputs/<paper_id>/report.md: final report
Do not commit PDFs, generated prompts, reviewer JSON, logs, editor bundles, reports, credentials, or local Codex runtime state. See SECURITY.md and docs/public_release_checklist.md.
Useful contributions include deterministic preprocessing improvements, reviewer prompts and validators, conservative normalization, tests, and cross-platform documentation. Use synthetic fixtures, public-domain examples, or short non-sensitive snippets in issues and pull requests.
Before sharing changes, run:
python -m unittest
python scripts/check_shareable_repo.py --include-untracked
python scripts/check_tracked_sensitive_names.py
git diff --checkSee CONTRIBUTING.md and docs/extension_guide.md.
For optional, source-adjudicated model comparisons, see the benchmark exercise. Preparation and scoring are offline; model runs require explicit opt-in.
MIT License. See LICENSE.md.