Repository navigation
GraphPaper 0.3: Science research mode, APA manuscripts and model-aware reasoning - #4
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cdf3daf8ec
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if p.research.plan.article_type=='empirical': | ||
| allowed={s.id for s in p.sources if s.enabled and s.role=='evidence'} | ||
| if not p.research.plan.empirical_results_source_ids or not set(p.research.plan.empirical_results_source_ids)<=allowed: | ||
| raise ValueError('An empirical article requires your own completed methods/results source. Select its source ID in the research plan; literature is not a new dataset.') |
There was a problem hiding this comment.
Reject scholarly papers as empirical result inputs
Require empirical result sources to be author-supplied and unlinked to retrieved research records, rather than accepting every enabled evidence source. As written, selecting an included paper's S# satisfies this guard, and draft_science() then passes that published paper to the model as author_empirical_material, allowing a literature source to be presented as the author's completed experiment.
Useful? React with 👍 / 👎.
| if plan.article_type in {'systematic_review','scoping_review'}: | ||
| if truncated:blockers.append('Systematic/scoping coverage is incomplete: at least one search is truncated. Narrow the documented query or complete retrieval outside GraphPaper and record it before claiming completeness.') | ||
| if any(r.decision=='unscreened' for r in p.research.records):blockers.append('Unscreened records remain.') | ||
| if not plan.inclusion or not plan.exclusion:blockers.append('Document explicit eligibility criteria.') | ||
| if any(s.get('status')!='ok' for s in p.research.searches):blockers.append('Complete or resolve failed databases before calling this a systematic search.') |
There was a problem hiding this comment.
Evaluate searches against the current systematic protocol
Validate the latest execution for every database/query in the current plan instead of applying any(... != 'ok') to the entire historical log. A transient failure remains forever after a successful rerun and permanently blocks submission, while changing the protocol to add a database or query can still pass using unrelated old successful logs, so the same check can both reject completed work and approve incomplete systematic coverage.
Useful? React with 👍 / 👎.
| pack=evidence_pack(p,min(45000,clients.settings.context_chars//2)) | ||
| raw=clients.complete(SCIENTIFIC,dump({'task':'Outline an APA scientific manuscript. Return the specified sections in order; no references or abstract section. Each section must advance a clear scientific question. For reviews, Method describes only actual logged searches and screening, Results is evidence synthesis rather than fabricated experimental data. For a protocol write planned procedures in future tense.', | ||
| 'plan':plan.model_dump(),'brief':p.brief.model_dump(),'exploratory_framing':next((a.model_dump() for a in p.angles if a.id==p.selected_angle),None),'required_sections':labels,'evidence':pack, | ||
| 'search_log':p.research.searches,'schema':{'title':'Specific manuscript title','sections':[{'title':'Introduction','purpose':'Scientific purpose','beats':['Specific point'],'source_ids':['S1'],'target_words':600}]}}),json_mode=True) |
There was a problem hiding this comment.
Refuse outlines that omit included evidence
Check that evidence_pack() returned one row for every included record before requesting an outline, as the drafting path already does. When a large evidence set exceeds the 45,000-character budget, evidence_pack() stops at the first row that no longer fits, but this path silently generates and accepts an outline based only on the leading subset of included studies.
Useful? React with 👍 / 👎.
| for query in queries: | ||
| for db in plan.databases: | ||
| job.check();job.note(f'Searching {NAMES[db]} ({complete+1}/{total})',5+int(80*complete/total)) | ||
| selected=plan.database_queries.get(db) or query | ||
| try: | ||
| rows,log=net.search(db,selected,plan) |
There was a problem hiding this comment.
Run each database-specific override only once
Move database overrides outside the loop over general queries, or define one override per query. With multiple planned queries and a single override for a database, selected resolves to the same override on every iteration, causing identical HTTP searches and duplicate audit-log entries while consuming rate limits and reporting them as separate completed searches.
Useful? React with 👍 / 👎.
Requested changes
Final validation on native Windows
Verified Windows artifact
GraphPaper-v0.3.0-Windows-x64.zip: 157,726,720 bytes.SHA-256: b32cb385726f185ee50c1347a31eb4631d924e7fc4247b46afc2003c14de0a1b.
Reports are under
docs/validation/v0.3.0; screenshots are actual UI captures with explicitly synthetic fixture material. The 139-file publishing manifest was verified against staged Git blobs. No user manuscripts, credentials, build environments or runtime binaries are committed. The Windows release pipeline requires the compiled native regression and will not silently overwrite an existing verified release asset.