Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -78,11 +78,11 @@ inventory-only row for the same provider id.

## Product surfaces

- CLI: `external-evidence discover|plan|receipt|admit|retire`.
- Managed Turn: the same five effect-runtime methods.
- Frontend/Lark: not changed in this Core slice. A companion slice should render
the same typed plan/admission projection and readback; it must not invent a
second registry or lifecycle.
- CLI: `external-evidence discover|plan|execute|receipt|admit|readback|retire`.
- Managed Turn: the same five typed effect-runtime methods; explicit provider
execution and ledger projection use the capability's CLI owner.
- Frontend/Lark: existing conversation answer/report and Markdown transports
render the shared validated readback. No independent registry or lifecycle.

## Acceptance

Expand All @@ -102,6 +102,29 @@ inventory-only row for the same provider id.
- retirement waits for downstream coverage of every admitted source;
- CLI and effect-runtime TypeScript tests pass from the source checkout.

## Delivery checkpoint (2026-10-02)

The public GitHub method now completes a bounded real journey: anonymous pinned
file reads, exact-plan receipt validation, a separate parent decision, projection
into the existing deepresearch source ledger, actual lineage readback and retirement.
Optional source refs and literal search terms are bound into the request/plan digest;
legacy requests retain their existing identity. The provider is bundled in extensions
under `method:public-github`; capability and ledger owners remain unchanged.

Passed: real public-provider/source CLI journey; negative cases for private or stale
readiness, malformed/unpinned sources, plan/admission mutation, partial/empty/failed
reads, independent admission and coverage, wrong-question projection, budget failure
and idempotent replay; packaged desktop/mobile conversation readback and reload;
existing Lark Markdown presentation. Source bodies are not persisted. The shared
Markdown readback uses existing answer/report and Lark transports; no frontend
configuration or parallel evidence authority is needed.

Commands are in the [versioned capability guide](../../../loopx/capabilities/external_research/README.md#public-github-method--公开-github-方法).
Live Lark delivery, authenticated connector execution and broader semantic research
quality remain untested; this checkpoint does not promote those providers or close
S6/S8. Failed or partial results preserve original-source fallback, and neither a
successful read nor a parent admission certifies evidence completeness.

## Non-goals

- a universal browser/search engine;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -60,10 +60,11 @@ Connector registry 继续只拥有库存与遥测。`supported` 绝不映射为

## 产品入口

- CLI:`external-evidence discover|plan|receipt|admit|retire`;
- Managed Turn:复用同五个 effect-runtime 方法;
- Frontend/Lark:本 Core 切片不修改。后续 companion slice 只渲染同源 plan/admission
投影与读回,不建立第二个 registry 或生命周期。
- CLI:`external-evidence discover|plan|execute|receipt|admit|readback|retire`;
- Managed Turn:复用五个 typed effect-runtime 方法;显式 provider 执行与账本投影
使用能力的 CLI owner;
- Frontend/Lark:现有会话答复/报告和 Markdown 运输渲染同源校验回读,
不建立独立 registry 或生命周期。

## 验收

Expand All @@ -80,6 +81,24 @@ Connector registry 继续只拥有库存与遥测。`supported` 绝不映射为
- 全部被采纳来源完成下游覆盖前不得退休;
- CLI 与 effect-runtime TypeScript 测试在源码 checkout 中通过。

## 交付检查点(2026-10-02)

公开 GitHub method 已完成有界真实链路:匿名读取固定提交文件、精确 plan 回执校验、
独立父 Agent 决定、投影到现有 deepresearch 来源账本、实际 lineage 回读与退休。
可选 source refs 和字面检索词进入 request/plan digest;旧请求身份保持兼容。
provider 以 `method:public-github` 内置在 extensions,能力和账本 owner 不变。

通过:真实公开 provider/源码 CLI 链路;私有或过期 readiness、无效/未固定来源、
plan/admission 篡改、部分/空/失败读取、独立采纳与覆盖、问题不匹配、预算耗尽及
幂等重放等负向用例;打包桌面/移动会话回读与重载;现有 Lark Markdown 展示。
不持久化来源正文。同源 Markdown 沿用现有答复/报告和 Lark 运输路径,
无需新增前端配置或并行证据权威。

命令参见[版本化能力指南](../../../loopx/capabilities/external_research/README.md#public-github-method--公开-github-方法)。
真实 Lark 送达、带凭据 connector 执行和更广泛语义研究质量尚未验证;
该检查点不晋升这些 provider,也不关闭 S6/S8。失败或部分结果保留原始来源退路;
读取成功和父 Agent 采纳均不证明证据完整性。

## 非目标

- 通用浏览器或搜索引擎;
Expand Down
2 changes: 2 additions & 0 deletions examples/personal-workspace-browser-smoke.mjs
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
#!/usr/bin/env node
import {nativeChildActivityScenario} from "./personal-workspace-browser/native-child-activity.mjs";
import {conversationImageRequestScenario} from "./personal-workspace-browser/conversation-image-request.mjs";
import {externalEvidenceReadbackScenario} from "./personal-workspace-browser/external-evidence-readback.mjs";
// Isolated browser acceptance scenarios for the personal Agent workspace.

import { mkdir, writeFile } from "node:fs/promises";
Expand Down Expand Up @@ -74,6 +75,7 @@ scenarioCatalog.push(goalWorkMapScenario);
scenarioCatalog.push(performanceDiagnosisScenario);
scenarioCatalog.push(blockedNoticeSettingsScenario);
scenarioCatalog.push(nativeChildActivityScenario);
scenarioCatalog.push(externalEvidenceReadbackScenario);
const requestedScenario = process.env.LOOPX_PERSONAL_WORKSPACE_SCENARIO;
const scenarios = requestedScenario
? scenarioCatalog.filter((scenario) => scenario.id === requestedScenario)
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
import assert from "node:assert/strict";
import { spawnSync } from "node:child_process";
import { resolve } from "node:path";
import { resolveTestPython } from "../../scripts/test-python.mjs";
import { outputDir, repoRoot } from "./fixture.mjs";
import { openWorkspacePage } from "./scenario-context.mjs";

export const externalEvidenceReadbackScenario = {
id: "external-evidence-readback",
async run({browser, collectCoverage, url}) {
// Exercise the product's typed plan/admission and actual downstream ledger,
// then hand the resulting shared readback to the existing answer surface.
const result = spawnSync(resolveTestPython(), ["-X", "utf8", "-c", `
import tempfile
from pathlib import Path
from loopx.control_plane.effect_runtime import effect_runtime_result as effect
from loopx.capabilities.deep_research.runtime import start_research
from loopx.capabilities.external_research.projection import readback, render_readback
provider = {"provider_id":"method:public-github", "provider_kind":"method", "protocol":"external_evidence_research_v0", "declared":True, "installed":True, "enabled":True, "ready":True, "unavailable_reason":None}
ref = "https://github.com/example/public/blob/" + "a"*40 + "/README.md"
plan = effect("external_evidence.plan", {"request":{"objective":"Inspect public fixture", "user_activity":"Choose a source", "decision":"Whether to use the fixture", "evidence_kinds":["literal_match"], "source_refs":[ref]}, "providers":[provider]})
receipt = {"schema_version":"loopx_external_evidence_receipt_v0", "plan_id":plan["plan_id"], "request_id":plan["request"]["request_id"], "provider_id":provider["provider_id"], "provider_kind":"method", "status":"succeeded", "summary":"Read pinned public fixture", "completed_at":"2026-10-02T00:00:00Z", "sources":[{"source_ref":ref, "source_family":"github_repository_file", "basis":"observed", "finding":"Literal fixture marker observed at line 1", "content_digest":"sha256:"+"b"*64, "accessed_at":"2026-10-02T00:00:00Z", "limitation":"Retrieval only; completeness unverified"}]}
admission = effect("external_evidence.admit", {"plan":plan, "receipt":receipt, "decision":{"disposition":"admit", "reason":"Synthetic parent checked the direct source", "admitted_source_refs":[ref]}})
with tempfile.TemporaryDirectory(prefix="lxe-ui-") as folder:
project=Path(folder)
start_research(project, question=plan["request"]["objective"], max_sources=8, max_subquestions=4)
print(render_readback(readback(plan, receipt, admission, project=project, execute=True)))
`], {cwd:repoRoot, encoding:"utf8", env:{...process.env, PYTHONPATH:repoRoot}, timeout:45000});
assert.equal(result.status, 0, result.stderr);
const markdown = result.stdout;
const context = await openWorkspacePage(browser, url, {collectCoverage});
const {api, page} = context;
try {
api.answerForMessage = () => markdown;
await page.getByRole("navigation", {name:"管家视图"}).getByRole("button", {name:/^(Chat|对话)$/}).click();
await page.getByLabel("向 LoopX 发送消息").fill("Show the external evidence readback");
await page.getByRole("button", {name:"发送",exact:true}).click();
const answer = page.locator(".personal-channel-timeline .personal-message.is-assistant", {hasText:"External evidence"});
await answer.waitFor();
assert.match(await answer.innerText(), /retire_ready/u);
assert.match(await answer.innerText(), /Downstream coverage.*1 admitted/u);
assert.match(await answer.innerText(), /Evidence completeness is unverified/u);
await answer.scrollIntoViewIfNeeded();
await page.screenshot({path:resolve(outputDir,"external-evidence-desktop.png"),fullPage:false,animations:"disabled"});
await page.setViewportSize({width:390,height:844});
await answer.scrollIntoViewIfNeeded();
assert.equal(await answer.evaluate(el => el.scrollWidth <= el.clientWidth + 1), true);
await page.screenshot({path:resolve(outputDir,"external-evidence-mobile.png"),fullPage:false,animations:"disabled"});
await page.reload({waitUntil:"networkidle"});
await page.getByRole("navigation", {name:"管家视图"}).getByRole("button", {name:/^(Chat|对话)$/}).click();
await answer.waitFor();
assert.match(await answer.innerText(), /retire_ready/u);
assert.equal(context.errors.length, 0);
return {coverageEntries:context.coverageEntries, note:"Shared typed evidence readback survives conversation reload and fits desktop/mobile."};
} finally { await context.close(); }
},
};
76 changes: 76 additions & 0 deletions examples/public-github-evidence-live-smoke.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
#!/usr/bin/env python3
"""Opt-in public GitHub exact-plan journey through the shipped source CLI.

Only anonymous public GETs and a disposable synthetic research ledger are used.
The explicit synthetic parent decision is separate from provider execution.
"""
from __future__ import annotations

import argparse
import json
import subprocess
import sys
import tempfile
from pathlib import Path


def qualify(source: str, root: Path) -> dict:
def run(*args, json_output=True):
result = subprocess.run([sys.executable, "-X", "utf8", "-m", "loopx.entrypoint",
"--runtime-root", str(root / "runtime"), "--registry", str(root / "registry.json"),
*args, "--format", "json" if json_output else "markdown"], capture_output=True,
text=True, encoding="utf-8", timeout=90, check=True)
return json.loads(result.stdout) if json_output else result.stdout

def save(name, value):
path = root / name
path.write_text(json.dumps(value), encoding="utf-8")
return str(path)

objective = "Inspect the public pinned README for the LoopX literal"
plan = run("external-evidence", "plan", "--objective", objective,
"--user-activity", "Choose a public source", "--decision", "Whether the source contains LoopX",
"--evidence-kind", "literal_match", "--public-github", "--source", source, "--search-term", "LoopX")
assert plan["status"] == "ready"
plan_path = save("plan.json", plan)
execution = run("external-evidence", "execute", "--plan-json", plan_path, "--execute")
receipt = execution["receipt"]
assert receipt["status"] == "succeeded" and len(receipt["sources"]) == 1
assert "observed at lines" in receipt["sources"][0]["finding"]
receipt_path = save("receipt.json", execution)
before = run("external-evidence", "readback", "--plan-json", plan_path, "--receipt-json", receipt_path)
assert before["parent_admission"] is None and before["retirement"] is None
admitted = run("external-evidence", "admit", "--plan-json", plan_path, "--receipt-json", receipt_path,
"--decision", "admit", "--reason", "Synthetic parent checked the pinned file and literal-match finding",
"--admit-source", source)
admission_path = save("admission.json", admitted)
project = root / "research"
run("deepresearch", "start", "--project", str(project), "--question", objective)
args = ("external-evidence", "readback", "--plan-json", plan_path, "--receipt-json", receipt_path,
"--admission-json", admission_path, "--project", str(project))
assert run(*args)["retirement"]["status"] == "retained"
result = run(*args, "--execute")
assert result["downstream_source_refs"] == [source]
assert result["retirement"]["status"] == "retire_ready"
assert run(*args, "--execute") == result
markdown = run(*args, json_output=False)
assert source in markdown and "retire_ready" in markdown and "Evidence completeness is unverified" in markdown
return {"ok": True, "provider": "method:public-github", "sources_observed": 1,
"explicit_parent_admission": True, "actual_ledger_readback": True,
"retained_before_projection": True, "idempotent_projection": True,
"raw_content_persisted": False, "source_ref": source}


def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--execute-public-provider", action="store_true")
parser.add_argument("--source", required=True)
args = parser.parse_args()
if not args.execute_public_provider:
parser.error("--execute-public-provider is required for anonymous public GETs")
with tempfile.TemporaryDirectory(prefix="lxe-") as folder:
print(json.dumps(qualify(args.source, Path(folder)), sort_keys=True))


if __name__ == "__main__":
main()
24 changes: 23 additions & 1 deletion loopx/capabilities/deep_research/runtime.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@
from pathlib import Path
from typing import Any

from ...control_plane.content_digest import ENVELOPED_SHA256_PATTERN
from ...file_lock import exclusive_file_lock

COMMAND = "/loopx-deepresearch"
Expand Down Expand Up @@ -238,20 +239,40 @@ def add_source(
tool: str,
title: str | None,
claims: list[dict[str, Any]],
external_evidence: dict[str, str] | None = None,
expected_question: str | None = None,
) -> dict[str, Any]:
# Load, validate the whole batch, allocate ids, mutate, and save all under
# one project-level lock: concurrent deepresearch commands are a normal
# agent-host shape, and an unlocked read-modify-write loses updates even
# when the atomic rename itself succeeds.
with exclusive_file_lock(state_path(project), operation="deepresearch_add_source"):
state = _require_active_state(project)
if expected_question is not None and state["question"] != expected_question:
raise ValueError("external evidence objective does not match the active research question")
if external_evidence is not None:
fields = {"plan_id", "admission_id", "receipt_digest", "content_digest"}
if set(external_evidence) != fields or any(
not isinstance(value, str) or not ENVELOPED_SHA256_PATTERN.fullmatch(value)
for value in external_evidence.values()
):
raise ValueError("external evidence lineage requires exact content-addressed identities")
url_or_path = url_or_path.strip()
tool = tool.strip() or "unspecified"
if not url_or_path:
raise ValueError("--url-or-path must be non-empty")
normalized = _normalize_source_ref(url_or_path)
for source in state["sources"]:
if _normalize_source_ref(str(source["url_or_path"])) == normalized:
same_source = (
str(source["url_or_path"]).strip().rstrip("/") == url_or_path.rstrip("/")
if external_evidence is not None else
_normalize_source_ref(str(source["url_or_path"])) == normalized
)
if same_source:
if (external_evidence is not None and source.get("external_evidence") == external_evidence
and [claim["text"] for claim in state["claims"] if claim["id"] in source["claims"]]
== [str(claim.get("text", "")).strip() for claim in claims]):
return {"source_id": source["id"], "claim_ids": source["claims"], "state": state}
raise ValueError(
f"source already recorded as {source['id']} "
f"({source['url_or_path']}); reuse its claims instead of re-reading"
Expand Down Expand Up @@ -325,6 +346,7 @@ def add_source(
"title": (title or "").strip() or None,
"accessed_at": _now_iso(),
"claims": claim_ids,
**({"external_evidence": dict(external_evidence)} if external_evidence is not None else {}),
}
)
_save_state(project, state)
Expand Down
Loading
Loading