Skip to content

GraphPaper 0.3: Science research mode, APA manuscripts and model-aware reasoning - #4

Merged
AronAxe merged 2 commits into
mainfrom
feat/science-reasoning-v0.3.0
Oct 7, 2026
Merged

AronAxe merged 2 commits into
mainfrom
feat/science-reasoning-v0.3.0

Conversation

@AronAxe

@AronAxe AronAxe commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Requested changes

  • Independent writer, editor and extraction/research reasoning levels. Codex supplies its actual advertised capabilities (including extended levels only when supported); OpenRouter, compatible APIs and Anthropic use provider-native parameters. Provider default sends no override and unsupported combinations are not silently downgraded.
  • Third Science workspace: live PubMed, Semantic Scholar, arXiv, Crossref and Europe PMC searches; query/status/count logs; DOI/identifier deduplication; human screening; accessible full text and uploaded-paper attachment; evidence appraisal and scientific outlines/drafting.
  • APA 7 professional Word manuscripts with metadata-derived author-year citations and hanging references. Research packages include BibTeX/RIS, the evidence matrix, protocol and actual search log. Empirical mode requires the author's completed results; submission checks do not fabricate ethics approval, exhaustive search coverage or peer-review acceptance.
  • Preserves author voice, project folders, humanizer/deslop, Codex OAuth and the v0.2.1 private native bridge/nonblocking close fixes.

Final validation on native Windows

  • 172 automated tests passed; one symlink-permission test skipped.
  • All 39 real-HTTP browser workflow checks passed, zero JavaScript page errors.
  • All 10 native source checks and all 10 checks against the actual compiled executable passed, including Science protocol editing, native Save As, the restricted bridge and pending-edit persistence before normal process exit.
  • Actual metadata retrieval succeeded on PubMed, arXiv, Crossref and Europe PMC. Anonymous Semantic Scholar returned HTTP 429, explicitly logged; its optional API-key route is implemented.
  • The same APA exporter source received a four-page rendered visual inspection. Blank-page and inherited font-theme issues found during QA were corrected.
  • Official Codex 0.160.1 signed-out runtime checks passed. Live paid model/JEV calls and a real user OAuth ceremony are not claimed as tested.
  • Added streaming ZIP packaging to avoid large in-memory compression buffers; CRC and SHA-256 validation passed.

Verified Windows artifact

GraphPaper-v0.3.0-Windows-x64.zip: 157,726,720 bytes.

SHA-256: b32cb385726f185ee50c1347a31eb4631d924e7fc4247b46afc2003c14de0a1b.

Reports are under docs/validation/v0.3.0; screenshots are actual UI captures with explicitly synthetic fixture material. The 139-file publishing manifest was verified against staged Git blobs. No user manuscripts, credentials, build environments or runtime binaries are committed. The Windows release pipeline requires the compiled native regression and will not silently overwrite an existing verified release asset.

@AronAxe
AronAxe marked this pull request as ready for review October 7, 2026 07:33
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T07:38:12.178561Z cdf3daf Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@AronAxe
AronAxe merged commit 967666c into main Oct 7, 2026
8 of 10 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cdf3daf8ec

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread graphpaper/science.py
Comment on lines +175 to +178
if p.research.plan.article_type=='empirical':
allowed={s.id for s in p.sources if s.enabled and s.role=='evidence'}
if not p.research.plan.empirical_results_source_ids or not set(p.research.plan.empirical_results_source_ids)<=allowed:
raise ValueError('An empirical article requires your own completed methods/results source. Select its source ID in the research plan; literature is not a new dataset.')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reject scholarly papers as empirical result inputs

Require empirical result sources to be author-supplied and unlinked to retrieved research records, rather than accepting every enabled evidence source. As written, selecting an included paper's S# satisfies this guard, and draft_science() then passes that published paper to the model as author_empirical_material, allowing a literature source to be presented as the author's completed experiment.

Useful? React with 👍 / 👎.

Comment thread graphpaper/science.py
Comment on lines +275 to +279
if plan.article_type in {'systematic_review','scoping_review'}:
if truncated:blockers.append('Systematic/scoping coverage is incomplete: at least one search is truncated. Narrow the documented query or complete retrieval outside GraphPaper and record it before claiming completeness.')
if any(r.decision=='unscreened' for r in p.research.records):blockers.append('Unscreened records remain.')
if not plan.inclusion or not plan.exclusion:blockers.append('Document explicit eligibility criteria.')
if any(s.get('status')!='ok' for s in p.research.searches):blockers.append('Complete or resolve failed databases before calling this a systematic search.')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Evaluate searches against the current systematic protocol

Validate the latest execution for every database/query in the current plan instead of applying any(... != 'ok') to the entire historical log. A transient failure remains forever after a successful rerun and permanently blocks submission, while changing the protocol to add a database or query can still pass using unrelated old successful logs, so the same check can both reject completed work and approve incomplete systematic coverage.

Useful? React with 👍 / 👎.

Comment thread graphpaper/science.py
Comment on lines +185 to +188
pack=evidence_pack(p,min(45000,clients.settings.context_chars//2))
raw=clients.complete(SCIENTIFIC,dump({'task':'Outline an APA scientific manuscript. Return the specified sections in order; no references or abstract section. Each section must advance a clear scientific question. For reviews, Method describes only actual logged searches and screening, Results is evidence synthesis rather than fabricated experimental data. For a protocol write planned procedures in future tense.',
'plan':plan.model_dump(),'brief':p.brief.model_dump(),'exploratory_framing':next((a.model_dump() for a in p.angles if a.id==p.selected_angle),None),'required_sections':labels,'evidence':pack,
'search_log':p.research.searches,'schema':{'title':'Specific manuscript title','sections':[{'title':'Introduction','purpose':'Scientific purpose','beats':['Specific point'],'source_ids':['S1'],'target_words':600}]}}),json_mode=True)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Refuse outlines that omit included evidence

Check that evidence_pack() returned one row for every included record before requesting an outline, as the drafting path already does. When a large evidence set exceeds the 45,000-character budget, evidence_pack() stops at the first row that no longer fits, but this path silently generates and accepts an outline based only on the leading subset of included studies.

Useful? React with 👍 / 👎.

Comment thread graphpaper/science.py
Comment on lines +69 to +74
for query in queries:
for db in plan.databases:
job.check();job.note(f'Searching {NAMES[db]} ({complete+1}/{total})',5+int(80*complete/total))
selected=plan.database_queries.get(db) or query
try:
rows,log=net.search(db,selected,plan)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Run each database-specific override only once

Move database overrides outside the loop over general queries, or define one override per query. With multiple planned queries and a single override for a database, selected resolves to the same override on every iteration, causing identical HTTP searches and duplicate audit-log entries while consuming rate limits and reporting them as separate completed searches.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant