One existing, tool-capable agent controls a user-authorized workflow. The user does not manually coordinate several agents. The local core has no model SDK. A Skill describes the domain procedure; the model uses normal shell/file tools.
The scanner does factual extraction. The host agent inspects source and history, then writes semantic patches. The core validates these patches and persists transactions. The frontend is a read-only projection, not the source of truth.
A research workspace owns a generated project ID saved in manifest.json.
It is not derived from a directory, Git remote spelling, or branch. Sources have
explicit aliases and machine-local paths. Runs prefer recorded run_id; otherwise
source alias plus relative run directory is a clearly labeled fallback.
A folder rename with no recorded identity needs reconciliation; identical config or hash is not enough to merge scientific identity. Changed configuration attached to an existing cell is flagged and excluded rather than silently reassigned. New independent seeds are distinct runs. A shared baseline is one set of runs referenced by multiple research questions.
The authoritative state is the ordered event directory. Each event has revision, previous hash, timestamp, actor, reason and upserted records. Writes use a local process lock and an expected base revision. An event is flushed to a temporary file and atomically renamed into place. Reads replay and validate the event chain. No LLM regenerates the old state. Same content is a no-op. Failed validation does not partially update state.
This detects accidental corruption, but is not a cryptographic signature or append-only filesystem. An OS user can rewrite files and recompute hashes. It is not a distributed system; competing Git edits require deliberate reconciliation. Replay is linear in history size. Production work includes snapshots and schema migration, without making caches authoritative.
module, question, cell, run, evaluation, claim, decision, artifact
are graph objects. parent_id is the navigational tree. Typed edges capture
comparison, support, contradiction, derivation, checkpoint inheritance and other
explicit relationships. They do not encode automatically proved causality.
Evidence references bind a file ID to a specific content hash and locator. Accepted claims and decisions require a direct immutable source and a scope. Linked node IDs are convenient navigation, not independent proof. Later source changes flag applicability review without erasing the original snapshot.
A decision has required context, covered factor values, replication keys and reopening conditions. It only counts as approved reuse after explicit human confirmation. Unknown or extra conditions yield review, not automatic training. The CLI is not a security boundary against an agent with arbitrary shell access; host permissions and human authorization remain essential.
Core ingestion recognizes metrics.json, results.json, metrics.csv,
config.json, config.toml, and optional run.json. The first metric filename
by that priority is used per run directory; additional sources are inventoried
and flagged, not silently pooled. If the primary file disappears and another
becomes primary, the run enters needs_review and retains old evidence until
an agent reconciles the transition. Conflicting protocol fields across metric,
run and configuration sources also block pooling. CSV uses the last row and does
not invent completion. JSON requires explicit completion to enter summaries.
Large binary files are metadata only; no pickle/torch load. Known secret names, symlinks and dependency/internal directories are excluded. Source code is parsed with Python AST only. YAML and arbitrary files remain evidence for host inspection. There is no execution of instructions in input files.
Coverage is layered: directory inventory, selected and verified source bytes, indexed files, agent-mapped graph objects, and independently replayed experiments are separate counts. Selected files do not equal independent experiments. When a project has corrected results or later branches, the agent records their identities and completion scopes separately with source evidence.
The default Research tree is one SVG canvas with horizontal lineage depth and vertical research tracks. Candidates are compact colored points with short ids; full labels appear on hover and details appear only on selection. Metrics and source records stay in the inspector, with runs/evidence behind a disclosure. Outline remains available for the original full-text hierarchy.
The layout groups by the nearest module/question and uses explicit derived_from
edges from origin to derivative, including cross-track parentage. Candidate
hierarchy is a fallback only when no explicit parent was recorded. A stable
primary-parent projection drives depth and folding; secondary parents, inspirations
and cycle-closing edges are retained and can be toggled on. Missing parents and
cycles remain warnings. L0/L1/etc denote computed lineage depth, not experimental
timestamps or claimed native generations. Colors use recorded selection/failure
information, not an inferred scientific verdict. Pan, zoom, fit, expanded canvas,
search and descendant folding modify presentation only; all identities, evidence
and original relations are preserved. The public Simate RoboScientist evolution
canvas informed the visual design; no external source or assets are bundled.
Bundled HTML/CSS/JS use no CDN, frontend build or runtime dependency. The local HTTP service exposes only state and known text evidence. It binds to loopback, checks Host, requires a random API token and rejects cross-site fetches. There is no shell, file-write, experiment-launch or synchronization HTTP endpoint.
Offline export embeds one state snapshot. Source text is omitted by default;
--include-evidence opts in to embedding indexed text snapshots, including
indexed files not linked to semantic nodes, subject to a combined 5 MiB limit.
The offline source list marks omitted text and disables opening it. JSON
escapes <, >, & and Unicode separators; the frontend escapes dynamic labels
and uses textContent for evidence. Exports are still private artifacts and must
be reviewed before public sharing.
export-site writes one standalone snapshot per workspace plus an index.html
entry point. The top project switcher links between pages; it does not combine
their states, evidence, Cell comparisons or histories. The first workspace is
the site's default page. Evidence inclusion remains an explicit choice for the
entire private site.
The static 中文 / English switch defaults to Chinese and persists a local
preference. Nodes may optionally provide translated labels and descriptions in
attrs.i18n.zh / attrs.i18n.en; absent translations use the recorded original.
The switch never translates source evidence or mutates events. The research
timeline shows agent-recorded project events only when they have an explicit
source date and direct evidence; it distinguishes occurrence, period start and
report as-of dates. The separate record-revisions tab derives file, node and
edge changes from each LabWeft event. Node changes keep
their module assignment at the event revision; file-module filtering uses current
evidence references. Event timestamps record LabWeft observations or writes,
not historical experiment times. A node's archived status is a LabWeft record
state, not a claim that source files or an external project archive were changed.
An archived Cell is absent from the current map hierarchy and numeric comparison;
archived claims and decisions remain in the historical decisions list with their
evidence. Archiving does not delete their source snapshots.
Recorded Run-to-Run derivations are projected between their owning Cells for display. The projection preserves source Run edges, merges duplicate visual links, and omits intra-Cell replication edges. It does not assert common ancestry for every replicate or alter identities. Missing origins, multiple parents and projection cycles remain reviewable; no lineage is inferred from labels or dates.
A project-level module may name a stable attrs.current_focus node. The default
canvas then shows that line and its recorded origins in one lane. All research
tracks and Outline retain access to the complete history. Lane labels preserve
the recorded research topic even when every condition is a root.
A Run focus resolves to its owning Cell. Archived focus declarations are ignored;
unsupported focus targets are reported without hiding the overview. Hierarchy
ancestors are retained when they provide the recorded lineage fallback.
organization.py provides read-only structural hints in context, validate
and presentation data. A large condition list without synthesis or comparison,
wide peer sections and invalid focus references request review. Independent
study designs may legitimately have no lineage. These hints neither invalidate
schema nor certify scientific meaning; the host still authors the research
reasoning separately from implementation flow and the run ledger.
The optional autoresearch.py workflow and backends/shinka.py adapter execute
only through explicit CLI commands. Small task inputs are snapshotted and hashed.
A machine-local worker lease serializes each activity; a control lock orders
batch reservations and stop requests. Shinka owns candidate search and local job
dispatch. Bounded native runs provide drained pause boundaries, retaining the
same database and absolute generation budget. Ambiguous crashes do not auto-resume.
The importer reads native SQLite reports without running source or importing SDKs, then writes deterministic Cell/Run/Evaluation identities through ordinary Store transactions and immutable evidence blobs. Search scores are backend attributes, not completed training-seed metrics or accepted scientific claims. See the execution contract and limitations. The scanner, original semantic protocol and read-only HTTP service do not gain execution privileges.
Skills for Codex and Claude use their respective discovery locations. Generic agents can read the same Skill and command protocol. This is capability-based integration: the host must read files and run commands. It is not an API router, model fallback service, OAuth proxy or universal compatibility guarantee.