English · 中文
Academic literature discovery as a Skill.
Built natively for Claude Code; runs in Codex and any agent that loads the SKILL.md format.
Five open sources + native-Chinese search + AI-venue acceptances · four tiers · journal partitions · paste-ready search strategies for WOS / Scopus / Embase / 知网 · single-file Shadcn report.
→ Live demo (opens the actual report in your browser)
If it looks useful, a star makes it easy to find later and helps it reach the next person who needs it.
You ask your agent for papers; this Skill runs a real multi-source literature search across OpenAlex · Semantic Scholar · CrossRef · PubMed · arXiv — plus native-Chinese sources NSSD (社会科学) and yiigle (中华医学) when you query in Chinese, and OpenReview (ICLR / NeurIPS / ICML acceptances) for AI topics — classifies relevance via parallel LLM SubAgents, and writes a self-contained HTML report you can open in any browser. No external LLM keys — your agent is the LLM.
For the databases it can't self-containedly reach (Web of Science, Scopus, Embase, 知网 CNKI, 万方, SinoMed …), it does the next best thing: on an Audit run or when you ask ("给我 WOS 检索式" / "我要去知网查"), it writes you a paste-ready professional search strategy per platform — correct field tags, controlled vocabulary + free-text double-track, per-host syntax — distilled from the Cochrane Handbook / PRESS / PRISMA-S, with each strategy carrying a three-state verification label and its review points. It's a professional first draft with flagged review points, not a sign-off-ready deliverable — and it never scrapes a closed database.
In your agent's chat, after install:
Find papers on working memory training in older adults
CJK in the query routes to Chinese UI; otherwise English. Both report flavors use the same data pipeline.
| Scenario | Tier | What you get |
|---|---|---|
| Scope a topic before writing the proposal | Quick / Standard | 20–60 top-RCS papers + 300-word executive summary in 8 min |
| Background section for course paper / thesis chapter | Standard | 60–180 papers screened, BibTeX ready for Zotero / Mendeley |
| Writing a review article — need real domain coverage | Deep | 180–400 papers + 1-hop citation chasing + topic clustering |
| SR-prep: PRISMA log + reproducibility audit | Audit | 400–1000+ papers, PRISMA-S 16-item disclosure, MeSH-precise |
| Onboarding a research assistant in a new field | Quick | Hand them the report.html — three tabs, hover for context |
Clone into your agent's Skills directory. The target path varies by agent — pick the one matching what you use:
# Pick ONE target directory, depending on your agent:
# ~/.claude/skills/paper-search-pro # Claude Code
# ~/.codex/skills/paper-search-pro # Codex CLI
# ~/.agents/skills/paper-search-pro # cross-agent convention (Goose, Roo, etc.)
# ~/.config/opencode/skills/paper-search-pro # OpenCode
# ~/.codeium/windsurf/skills/paper-search-pro # Windsurf
# ./.claude/skills/paper-search-pro # project-local (Cursor, etc.)
git clone https://github.com/O0000-code/paper-search-pro.git \
~/.claude/skills/paper-search-pro # change to your chosen path
# Point PSP_HOME at wherever you cloned. SKILL.md STEP 0 also auto-resolves
# the common paths above, so this manual export is only needed for non-standard locations.
export PSP_HOME="$HOME/.claude/skills/paper-search-pro"
python3 -m pip install -r "$PSP_HOME/scripts/requirements.txt"
# Optional: install the exact dependency versions validated by CI.
python3 -m pip install -r "$PSP_HOME/scripts/requirements.lock"Five free API keys (~15 min total) — see references/setup.md.
The Skill picks Standard by default. Wording like thorough, systematic review, or 几篇 overrides.
| Tier | Wall-clock | Papers | Trigger signals | |
|---|---|---|---|---|
▏ |
Quick | 5–8 min | 20–60 | scan · "5 papers" · before-tomorrow scope |
▍ |
Standard | 10–17 min | 60–180 | default — background reading · course paper |
▋ |
Deep | 30–45 min | 180–400 | review article · thorough coverage · 综述写作 |
█ |
Audit | 2–3 hr | 400–1000+ | systematic review · PRISMA · Cochrane-adjacent |
Three tabs · three hero layouts · two list densities · responsive at 860 px · bilingual UI · Noto Sans SC inlined · fully offline.
Each report has a scholarly, evidence-bounded title authored after screening. The verbatim user request remains in audit metadata, while language, date, database, journal-tier, and export constraints stay in Methods rather than being repeated in the visual H1.
Journal partitions. Every paper is tagged with its 中科院 (CAS) / JCR / SJR tier — a quiet badge on each card, the full three-platform breakdown in the detail panel, and a zone filter (Q1 / ≥Q2 / ≥Q3, or 一区 / ≥二区) in the toolbar. The tables are fetched at runtime from public mirrors — never bundled — and attributed; only JCR's figure is labelled an impact factor.
Outputs land in $PWD/paper-search-results/<search_id>/:
report.html Self-contained Shadcn report (opens directly in browser)
report.md Markdown variant for pandoc · citation managers
papers.csv Spreadsheet export
papers.bib BibTeX for Zotero · Mendeley · LaTeX
papers.ris RIS for EndNote · Papers
papers.json Full structured data (UnifiedPaperEntity[])
kg_classified.json Internal KG with per-paper RCS scores
execution_log.json PRISMA-S 16-item disclosure log
summary.md 300-word executive summary in the main agent's voice
metadata.json Separate original request · search topic · display title
Five keys, all free, ~15 min total. Real config lives at ~/.paper-search-pro/config.yaml (mode 0600, auto-created). Template lives at assets/default_config.yaml.
| Layer | Source | Role | Cost | Apply at |
|---|---|---|---|---|
| L1 | OpenAlex | primary — always on | free | https://openalex.org/settings/api |
| L2 | PubMed | medical · MeSH enricher | free | https://account.ncbi.nlm.nih.gov/settings/ |
| L2 | arXiv | preprint freshness (T‑0~T‑4) · AI topics (T‑0~T‑365) | free | (no signup — SDK enforces 1 req / 3 s) |
| L3 | Semantic Scholar | influentialCitationCount + abstract fallback | free | https://www.semanticscholar.org/product/api |
| L3 | CrossRef | funder · license · clinical-trial-number | free | (no key — crossref_email only) |
Verify readiness any time:
PYTHONPATH=$PSP_HOME python3 -c \
"from scripts.config import load_config; c = load_config(); \
print('ready' if c.openalex_api_key and c.ncbi_email else 'missing')"Primary source — by preference, and by fallback. Two independent capabilities from one switch:
- Pick your primary. OpenAlex is the default; set Semantic Scholar as the primary source instead when its corpus or field coverage suits your topic better.
- Automatic quota fallback. When the active primary runs low on its daily quota (or starts erroring), the run continues on the other source rather than stopping — a depleted key degrades gracefully.
Beyond the five general sources in the table above, these three switch on only for their topics and need no key; each is run by an authority in its field:
| Source | Run by | When it is used | What it adds |
|---|---|---|---|
| NSSD (National Center for Philosophy and Social Sciences Documentation) | led by the Chinese Academy of Social Sciences | Chinese-language social-science and humanities queries | CSSCI and other Chinese social-science journals, which OpenAlex barely covers |
| yiigle (Chinese Medical Journals Full-text Database) | Chinese Medical Association Publishing House | Chinese-language medical queries | Chinese originals and abstracts of the Chinese Medical Association journals |
| OpenReview | non-profit platform run by a team at UMass Amherst | AI / machine-learning topics | Acceptance labels from ICLR, NeurIPS, ICML and others (e.g. "ICLR 2026 Oral"); accepted papers only |
A 14-step recipe in SKILL.md drives every run. Python helpers in scripts/ do deterministic API work; relevance classification is delegated to up to five parallel SubAgents per round. No third-party LLM keys required — your agent is the LLM.
┌──────────────────────────────────────────────────────────┐
│ Main agent reads SKILL.md (recipe) and drives the run │
└────────┬───────────────────────────────────────┬─────────┘
│ │
▼ ▼
┌───────────────┐ ┌───────────────┐
│ Python │ ← deterministic API │ SubAgents │
│ helpers │ work; no LLM, │ (parallel, │
│ │ no LLM key │ 5 / round) │
└───────┬───────┘ └───────┬───────┘
│ │
└───────────────────┬───────────────────┘
▼
┌─────────────────────────────┐
│ Single-file HTML report │
│ + MD + BibTeX + RIS │
│ + CSV + PRISMA-S log │
└─────────────────────────────┘
Per-step reference documents live in references/ — tier decisions, query planning (PICO / SPIDER / PEO), source routing, helper cheatsheets, the RCS rubric, stop conditions, citation chasing, the classifier SubAgent prompt, the PRISMA-S 16-item checklist, summary writer guide, error handling, output conventions.
Headless mode for agents. When the consumer is another agent rather than a person, scripts/agent_search.py runs the deterministic core in a single command and returns one structured JSON envelope — deduped, relevance-scored, saturation-checked — with no HTML and no classification SubAgents. Same search discipline, machine-readable output.
I maintain paper-search-pro mostly on my own, with help from two lovely contributors, @haoxinC111 and @MatrixA. You're welcome to be the third. If it has helped you, your star means a lot to me, and I'll keep improving it.
Apache License 2.0 — see LICENSE.txt.




