Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -135,6 +135,29 @@ An archived Cell is absent from the current map hierarchy and numeric comparison
archived claims and decisions remain in the historical decisions list with their
evidence. Archiving does not delete their source snapshots.

### Research organization review and focused lineage

Recorded Run-to-Run derivations are projected between their owning Cells for
display. The projection preserves source Run edges, merges duplicate visual
links, and omits intra-Cell replication edges. It does not assert common ancestry
for every replicate or alter identities. Missing origins, multiple parents and
projection cycles remain reviewable; no lineage is inferred from labels or dates.

A project-level module may name a stable `attrs.current_focus` node. The default
canvas then shows that line and its recorded origins in one lane. All research
tracks and Outline retain access to the complete history. Lane labels preserve
the recorded research topic even when every condition is a root.
A Run focus resolves to its owning Cell. Archived focus declarations are ignored;
unsupported focus targets are reported without hiding the overview. Hierarchy
ancestors are retained when they provide the recorded lineage fallback.

`organization.py` provides read-only structural hints in `context`, `validate`
and presentation data. A large condition list without synthesis or comparison,
wide peer sections and invalid focus references request review. Independent
study designs may legitimately have no lineage. These hints neither invalidate
schema nor certify scientific meaning; the host still authors the research
reasoning separately from implementation flow and the run ledger.

## Provider boundary

The optional `autoresearch.py` workflow and `backends/shinka.py` adapter execute
Expand Down
5 changes: 3 additions & 2 deletions research_harness/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
from .store import Store, HarnessError, atomic_write, read_json, json_write, validate_state
from .scan import sync
from .reasoning import compare, assess, drift
from .organization import organization
from .server import export_html, export_site, serve


Expand Down Expand Up @@ -222,7 +223,7 @@ def main(argv: list[str] | None = None) -> int:
top = [n for n in nodes if n["kind"] not in {"run", "evaluation", "artifact"}]
emit({"project": state["project"], "base_revision": state["revision"], "nodes": top,
"node_count": len(nodes), "file_count": len(state["files"]), "drift": drift(state),
"scan": state["scan"], "history": state["history"][-12:],
"scan": state["scan"], "history": state["history"][-12:], "organization": organization(state),
"next": "Use files for paginated evidence, context --node ID, or --full. Read evidence before applying a semantic patch."})
elif args.command == "files":
state = store.load()
Expand Down Expand Up @@ -253,7 +254,7 @@ def main(argv: list[str] | None = None) -> int:
except HarnessError as exc:
failures.append({"node": n["id"], "error": str(exc)})
emit({"valid": not failures, "revision": state["revision"], "nodes": len(state["nodes"]),
"evidence_errors": failures, "drift": drift(state),
"evidence_errors": failures, "drift": drift(state), "organization": organization(state),
"scope": "Mechanical validation only; not validation of scientific truth or host-model reliability."})
return 1 if failures else 0
elif args.command == "export":
Expand Down
70 changes: 70 additions & 0 deletions research_harness/organization.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
"""Read-only structural review hints; not a scientific-quality certificate."""
from __future__ import annotations


def organization(state: dict) -> dict:
nodes = {key: n for key, n in state['nodes'].items() if n.get('status') != 'archived'}
warnings = []
children = {}
for n in nodes.values():
children.setdefault(n.get('parent_id'), []).append(n)

def ancestor(node_id, kinds):
seen = set()
while node_id in nodes and node_id not in seen:
seen.add(node_id)
n = nodes[node_id]
if n['kind'] in kinds:
return node_id
node_id = n.get('parent_id')
return None

studies = {}
for n in nodes.values():
if n['kind'] == 'cell':
owner = ancestor(n.get('parent_id'), {'module', 'question'})
studies.setdefault(owner, []).append(n['id'])
related = set()
run_lineage = 0
for e in state.get('edges', {}).values():
if e['relation'] not in {'derived_from', 'compares', 'inspired_by'}:
continue
source, target = nodes.get(e['source']), nodes.get(e['target'])
if not source or not target or source['kind'] not in {'cell', 'run'} or target['kind'] not in {'cell', 'run'}:
continue
a, b = ancestor(source['id'], {'cell'}), ancestor(target['id'], {'cell'})
if a and b and a != b:
related.update((a, b))
if e['relation'] == 'derived_from' and (source['kind'] == 'run' or target['kind'] == 'run'):
run_lineage += 1
for owner, cells in studies.items():
if len(cells) < 5 or nodes.get(owner, {}).get('attrs', {}).get('role') == 'evidence_audit':
continue
attrs = nodes.get(owner, {}).get('attrs', {})
if not attrs.get('study_summary'):
warnings.append({'code': 'missing_study_synthesis', 'node_id': owner, 'cell_ids': cells,
'message': 'Several conditions have no study synthesis. Record the question, comparison, outcome and remaining gap on their study.'})
if not any(c in related for c in cells) and not attrs.get('design'):
warnings.append({'code': 'unexplained_condition_list', 'node_id': owner, 'cell_ids': cells,
'message': 'Conditions have neither recorded comparisons/lineage nor an explicit independent study design. Review evidence; never invent edges to make a tree.'})
for parent, items in children.items():
modules = [n for n in items if n['kind'] in {'module', 'question'} and n.get('attrs', {}).get('role') != 'evidence_audit']
if len(modules) >= 8:
warnings.append({'code': 'wide_research_outline', 'node_id': parent,
'related_ids': [n['id'] for n in modules],
'message': 'Many peer research sections. Review whether they form method families, comparisons, or historical diagnostics; retain independent topics when justified.'})
focuses = []
for n in nodes.values():
focus = n.get('attrs', {}).get('current_focus')
if focus is None:
continue
target = nodes.get(focus) if isinstance(focus, str) else None
if (not target or target['kind'] not in {'module', 'question', 'cell', 'run'}
or (target['kind'] == 'run' and ancestor(focus, {'cell'}) is None)):
warnings.append({'code': 'invalid_current_focus', 'node_id': n['id'],
'message': 'Current focus must reference a visible module, question, cell, or run with a visible owning cell.'})
else:
focuses.append(focus)
return {'needs_review': bool(warnings), 'warnings': warnings,
'run_lineage_edges': run_lineage, 'current_focus_ids': sorted(set(focuses)),
'scope': 'Structural review hints only. Independent experiments need not form a lineage; schema validity does not certify research organization or scientific conclusions.'}
3 changes: 2 additions & 1 deletion research_harness/server.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,13 +11,14 @@
import webbrowser
from .store import Store, HarnessError, atomic_write, canonical
from .reasoning import compare, drift
from .organization import organization

STATIC = Path(__file__).parent / "static"


def presentation(store: Store) -> dict:
state = store.load()
return {**state, "comparison": compare(state), "drift": drift(state), "mode": "live"}
return {**state, "comparison": compare(state), "drift": drift(state), "organization": organization(state), "mode": "live"}


def export_html(store: Store, output: Path, *, include_evidence: bool = False,
Expand Down
14 changes: 14 additions & 0 deletions research_harness/skill/research-harness/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,14 @@ below use legacy `rh` as shorthand for that same CLI. Quote all paths.
from sources, the researcher, or agent interpretation. Do not promote a
directory hierarchy into this outline. If accounts conflict, retain their
version and scope instead of choosing a convenient narrative.
Keep three separate explanations: research reasoning (failure/question ->
hypothesis -> controlled comparison -> result -> decision), the chosen
implementation's execution/data flow, and the full run ledger. A batch list
or runtime phase list is not a research reasoning chain. Explain why each
important result changed the next decision. Reuse the best-supported lineage
before opening another branch. When the user selects a main line, record its
stable ID as `attrs.current_focus` on the project-level research module;
retain parked studies and evidence as history.
5. **Map the experiment design and testable claims.** Use `rh context WORKSPACE
--full` or `--node ID` and reuse IDs. Nest modules for project-specific
research tracks, studies and ablation groups; connect questions → cells → runs
Expand Down Expand Up @@ -153,6 +161,12 @@ below use legacy `rh` as shorthand for that same CLI. Quote all paths.
human-protected corrections. Review source-drift flags without declaring every
old conclusion invalid.
9. **Check and visualize.** Run `rh validate WORKSPACE` and `rh compare WORKSPACE`.
Inspect `organization` in `context` and `validate`; `valid: true` checks
schema/evidence, not whether the research narrative is organized. Resolve or
explain structural review hints before delivery. Inspect the rendered map:
Run-level derivations must remain visible in the Cell overview. Never infer
parentage from version numbers or dates to make a prettier tree. A justified
independent comparison does not need artificial lineage edges.
Current comparisons exclude runs whose metric evidence changed or disappeared;
their old snapshots remain historical evidence. Address mechanical errors,
not by weakening provenance. Export with `rh export
Expand Down
23 changes: 23 additions & 0 deletions research_harness/skill/research-harness/references/PROTOCOL.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,29 @@ Relations: `derived_from`, `compares`, `supports`, `contradicts`,
`uses_checkpoint`, `supersedes`, `inspired_by`. Never confuse method derivation
with inherited weights. Every edge endpoint must exist.

## Research Organization and Display

The map projects recorded Run-to-Run derivations between their owning Cells.
This is read-only provenance aggregation, not a claim that every replicate has
the same ancestry. Duplicate visual links retain their Run sources; intra-Cell
replication is not a new condition. Multiple origins and cycles remain visible.
Record the actual Run relationship instead of inventing a Cell-level one for
the renderer.

An optional `attrs.current_focus: "STABLE_NODE_ID"` on a project-level module
selects the current research line. The focused view retains recorded origins;
All research tracks and Outline expose the full history. It does not change
scientific status, evidence, or run counts.
Use a visible module, question or Cell ID; a Run ID resolves to its owning Cell.
Archived declarations and unsupported focus targets do not hide the full map.

Keep the research reasoning chain separate from implementation execution flow
and the run ledger. Each major step identifies the question, baseline, change,
observed result and ensuing decision. Source names and runtime phases are not
research reasoning. The `organization` report in `context`/`validate` provides
structural review hints; warnings neither invalidate schema nor establish
scientific quality. Independent studies may legitimately have no lineage edges.

## Per-version change annotations

Every version, attempt or candidate shown in a research map needs a readable
Expand Down
4 changes: 4 additions & 0 deletions research_harness/static/app.css

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading
Loading