From 3cf771e99c2d4682c8f067916d78b7d925b20940 Mon Sep 17 00:00:00 2001 From: typakon4 Date: Mon, 21 Sep 2026 17:52:40 +0300 Subject: [PATCH 1/3] feat(hermes): add advisory Jev operator layer --- CHANGELOG.md | 2 + README.md | 6 +- README.ru.md | 6 +- README.zh-CN.md | 3 +- docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md | 177 +++++++++++++++++++++++++ docs/SCHEMA-VERSIONING.md | 4 +- integrations/hermes/__init__.py | 25 +++- integrations/hermes/plugin.yaml | 4 +- integrations/hermes/schemas.py | 26 ++++ package.json | 1 + plugin.yaml | 4 +- scripts/mcp-receipt-smoke.mjs | 15 +++ scripts/model-route-report.mjs | 15 +++ scripts/supervision-e2e.mjs | 31 ++++- src/hermes-adapter.mjs | 26 ++++ src/mcp-server.mjs | 54 ++++++++ src/model-route-metrics.mjs | 89 +++++++++++++ src/model-routing.mjs | 86 +++++++++++++ src/receipts.mjs | 1 + src/shadow-compaction.mjs | 178 ++++++++++++++++++++++++++ src/supervision.mjs | 34 ++++- test/hermes-adapter.test.mjs | 8 +- test/model-route-metrics.test.mjs | 61 +++++++++ test/model-routing.test.mjs | 59 +++++++++ test/shadow-compaction.test.mjs | 67 ++++++++++ test/supervision.test.mjs | 18 +++ 26 files changed, 987 insertions(+), 13 deletions(-) create mode 100644 docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md create mode 100644 scripts/model-route-report.mjs create mode 100644 src/model-route-metrics.mjs create mode 100644 src/model-routing.mjs create mode 100644 src/shadow-compaction.mjs create mode 100644 test/model-route-metrics.test.mjs create mode 100644 test/model-routing.test.mjs create mode 100644 test/shadow-compaction.test.mjs diff --git a/CHANGELOG.md b/CHANGELOG.md index 3927a44..3c18028 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -22,6 +22,8 @@ First public-release candidate. This version is prepared but has not been pushed - `jev_record_execution` and append-only JSONL routing/execution receipts joined by `correlation_id`. - Offline replay/evaluation without host execution calls. - Opt-in supervision judgments with deterministic host policy. +- Opt-in, report-only `jev_model_route` with correlated execution outcomes and a local evidence report. +- Hermes plugin integration with profile-scoped OpenRouter secret resolution, replay receipts, model-route correlation, and a local-agent handoff guide. - Opt-in deterministic context filtering (`shadow` and `conservative`). - Explicit capability discovery for skills, MCP, CLI, DSH, tools, subagents, and models. - Experimental, opt-in browser fast-path over host-supplied observations. diff --git a/README.md b/README.md index bbf59c9..46d2424 100644 --- a/README.md +++ b/README.md @@ -76,7 +76,9 @@ All three modes (`demo`, `openrouter`, and direct `typesafe`), their endpoints, - **Routing:** `jev_route` selects one capability from the host-supplied candidate set. Selection is advisory; the host validates the id and permissions. - **Receipts/replay:** `jev_record_execution` joins the host result to the original `correlation_id`. JSONL cases live in `.jev/replay/cases.jsonl` and can be evaluated offline with `npm run replay:evaluate`. -- **Supervision:** `jev_supervise` returns bounded work-state judgments; deterministic host policy maps them to `continue`, `verify`, `retry`, `finish`, or `escalate`. Jev does not perform those actions. +- **Supervision:** `jev_supervise` returns bounded work-state judgments; deterministic host policy maps them to `continue`, `verify`, `retry`, `finish`, or `escalate`. The receipt records a deterministic `evidence_state`: `present`, `missing`, or `contradictory`. Contradictory evidence cannot result in `finish`. Jev does not perform those actions. +- **Model routing:** `jev_model_route` returns one recommendation from host-declared model profiles for the next model call. It is shadow-only: the host must measure outcome, retries, latency, and cost before it changes any provider/model setting. Model-route decisions are joined to later `jev_record_execution` receipts by `correlation_id`; record `result.model_route = { actual_model_id, retry_count, outcome }`, then run `npm run model-route:report -- /path/to/cases.jsonl` for a read-only evidence report. +- **Shadow compaction:** `jev_shadow_compaction` makes batched, conservative, report-only keep/drop candidates for host-supplied context. It never mutates, summarizes, or deletes context; pinned paths, errors, commands, and requirements are retained without provider review, while provider failures retain every remaining item. - **Context filtering:** optional deterministic `shadow` or `conservative` filtering reduces stale context without LLM summarization. - **Experimental browser fast-path:** `jev_browser_step` chooses one bounded action from a host observation. The host supplies observations, approval, native execution, and recovery. It is opt-in and does not start a browser worker. - **Fail-open:** disabled, unavailable, invalid, or inconclusive Jev calls return control to the host's normal path. Jev never widens permissions or guesses execution. @@ -93,7 +95,7 @@ JEV_CONTEXT_FILTER=shadow jev cli --input examples/route-request.json Current examples live under `integrations/`: -- `integrations/hermes/` +- `integrations/hermes/` — see the [local-agent handoff (Russian)](docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md) for the active plugin topology and safe change workflow. - `integrations/omp/` - `integrations/codex/` - `integrations/template/` diff --git a/README.ru.md b/README.ru.md index 3abeb27..6ca87ea 100644 --- a/README.ru.md +++ b/README.ru.md @@ -76,7 +76,9 @@ jev doctor --project /path/to/workspace - **Routing:** `jev_route` выбирает одну capability из набора, предоставленного host. Host повторно проверяет id и permissions. - **Receipts/replay:** `jev_record_execution` связывает результат host с исходным `correlation_id`. JSONL-файлы находятся в `.jev/replay/cases.jsonl` и проверяются офлайн через `npm run replay:evaluate`. -- **Supervision:** `jev_supervise` возвращает ограниченные judgments о состоянии работы; детерминированная host policy преобразует их в `continue`, `verify`, `retry`, `finish` или `escalate`. Jev эти действия не выполняет. +- **Supervision:** `jev_supervise` возвращает ограниченные judgments о состоянии работы; детерминированная host policy преобразует их в `continue`, `verify`, `retry`, `finish` или `escalate`. Receipt хранит детерминированный `evidence_state`: `present`, `missing` или `contradictory`. Противоречивые evidence не могут привести к `finish`. Jev эти действия не выполняет. +- **Model routing:** `jev_model_route` возвращает одну рекомендацию из model profiles, которые объявил host, для следующего model call. Это только shadow: прежде чем менять provider/model setting, host обязан измерить outcome, retry, latency и cost. Решение связывается с последующим `jev_record_execution` по `correlation_id`; host записывает `result.model_route = { actual_model_id, retry_count, outcome }`, затем `npm run model-route:report -- /path/to/cases.jsonl` строит read-only evidence report. +- **Shadow compaction:** `jev_shadow_compaction` делает пакетный консервативный report с keep/drop-кандидатами для context, который передал host. Он никогда не меняет, не суммаризирует и не удаляет context; path, error, command и requirement pin-ятся до provider review, а provider failure оставляет все остальные items. - **Context filtering:** опциональная детерминированная фильтрация `shadow` или `conservative` убирает устаревший context без LLM-суммаризации. - **Experimental browser fast-path:** `jev_browser_step` выбирает одно ограниченное действие из observation host. Host предоставляет observation, approval, native execution и recovery. - **Fail-open:** при отключённом, недоступном, ошибочном или неубедительном Jev вызове управление возвращается в обычный host path. Jev не расширяет permissions и не угадывает execution. @@ -93,7 +95,7 @@ JEV_CONTEXT_FILTER=shadow jev cli --input examples/route-request.json Примеры находятся в `integrations/`: -- `integrations/hermes/` +- `integrations/hermes/` — для active plugin topology и safe change workflow см. [handoff локальному агенту (RU)](docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md). - `integrations/omp/` - `integrations/codex/` - `integrations/template/` diff --git a/README.zh-CN.md b/README.zh-CN.md index 643d452..2924467 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -76,7 +76,8 @@ jev doctor --project /path/to/workspace - **Routing:** `jev_route` 从 host 提供的候选集合中选择一个 capability。host 会再次验证 id 和权限。 - **Receipts/replay:** `jev_record_execution` 使用原始 `correlation_id` 关联 host 结果。JSONL 位于 `.jev/replay/cases.jsonl`,可用 `npm run replay:evaluate` 离线评估。 -- **Supervision:** `jev_supervise` 返回有界的工作状态判断;确定性的 host policy 将其映射为 `continue`、`verify`、`retry`、`finish` 或 `escalate`。Jev 不执行这些动作。 +- **Supervision:** `jev_supervise` 返回有界的工作状态判断;确定性的 host policy 将其映射为 `continue`、`verify`、`retry`、`finish` 或 `escalate`。receipt 记录确定性的 `evidence_state`:`present`、`missing` 或 `contradictory`;矛盾证据不能导致 `finish`。Jev 不执行这些动作。 +- **Shadow compaction:** `jev_shadow_compaction` 为 host 提供的 context 生成批量、保守、只报告的 keep/drop 候选。它不会修改、总结或删除 context;path、error、command 和 requirement 在 provider review 前被保留,provider 失败时保留所有其他 items。 - **Context filtering:** 可选的确定性 `shadow` 或 `conservative` 过滤器减少过期 context,不使用 LLM 摘要。 - **Experimental browser fast-path:** `jev_browser_step` 根据 host observation 选择一个有界浏览器动作。observation、审批、原生执行和恢复都由 host 提供。 - **Fail-open:** Jev 被禁用、不可用、出错或无法确定时,控制权返回 host 的正常路径。Jev 不扩大权限,也不猜测执行。 diff --git a/docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md b/docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md new file mode 100644 index 0000000..2ed5bc5 --- /dev/null +++ b/docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md @@ -0,0 +1,177 @@ +# Hermes integration: handoff для локального агента + +Этот документ описывает **текущую рабочую интеграцию** `jev-layer` в Hermes. Он нужен локальному агенту, чтобы сначала понять границы и точки входа, а не повторно «интегрировать Jev» через глобальный конфиг, произвольные shell-команды или второй executor. + +## Что уже установлено + +- Исходный репозиторий: `/home/hermes/jev-layer`. +- Активная plugin-копия Hermes: `~/.hermes/plugins/jev-layer`. +- Plugin регистрирует ровно шесть tools: + - `jev_route`; + - `jev_record_execution`; + - `jev_supervise`; + - `jev_browser_step`; + - `jev_shadow_compaction`; + - `jev_model_route`. +- Replay evidence активного профиля: `~/.hermes/jev-layer/replay/cases.jsonl`. +- Gateway загружает Python plugin при старте. После изменения `integrations/hermes/*`, `plugin.yaml` или списка tools нужен пользовательский `/restart` gateway. + +Не путай surfaces: + +```text +/home/hermes/jev-layer source checkout, тесты и документация +~/.hermes/plugins/jev-layer установленная plugin-копия текущего профиля +~/.hermes/jev-layer/replay/cases.jsonl profile-scoped evidence, не source artifact +~/.hermes/config.yaml глобальный Hermes config; не менять для этой plugin без явного задания +``` + +## Главный принцип + +```text +Hermes registry/permissions/approval/executor/recovery + │ closed, bounded input + ▼ + jev-layer decision + │ correlation_id, advice only + ▼ +Hermes validates → native execution → jev_record_execution → JSONL evidence +``` + +Jev **никогда** не получает право исполнять tool, shell, URL или browser action. Любое решение — advisory. Hermes остаётся владельцем discovery, policy, approval, native execution, retry, recovery и финального ответа. + +Если Jev выключен, недоступен, даёт невалидный ответ или низкую уверенность, текущий normal Hermes path продолжается. Не добавляй обходной executor и не меняй retry semantics. + +## Файловая карта + +| Зачем | Файл | +| --- | --- | +| Hermes tool registration + profile-scoped secret bridge | `integrations/hermes/__init__.py` | +| JSON schemas шести tool surfaces | `integrations/hermes/schemas.py` | +| Plugin manifest | `integrations/hermes/plugin.yaml`, `plugin.yaml` | +| Один JSONL request/response adapter process | `src/hermes-adapter.mjs` | +| Общая bounded routing contract | `src/route.mjs`, `src/contract.mjs` | +| Receipt/replay JSONL format | `src/receipts.mjs` | +| Browser action space, progress и recovery | `src/browser.mjs` | +| Model recommendation | `src/model-routing.mjs` | +| Model-route evidence report | `src/model-route-metrics.mjs`, `scripts/model-route-report.mjs` | +| Context keep/drop report | `src/shadow-compaction.mjs` | +| Work-state judgement | `src/supervision.mjs` | +| Harness-neutral integration contract | `docs/AGENT-IMPLEMENTATION.md` | + +## Как работают tools + +### `jev_route` + +Hermes сначала собирает **закрытый** `capabilities[]` из своего registry. Jev выбирает один известный id. До исполнения Hermes снова проверяет: `status`, точное совпадение id, availability, permission, risk и approval. После результата используется `jev_record_execution` с тем же `correlation_id`. + +Нельзя передавать в candidates hidden tools, raw shell command, credential, неограниченную историю или capability, для которой у Hermes нет native executor. + +### `jev_record_execution` + +Принимает результат уже выполненного/отклонённого host action. `correlation_id` должен принадлежать decision того же gateway process. Статусы: `completed`, `failed`, `not_started`. + +Receipt — evidence, не источник авторизации. В `result` нельзя писать токены, пароли, номера карт и полный неочищенный лог. + +### `jev_model_route` + +Возвращает рекомендацию одного из **host-declared** `models[]`. Всегда `route_mode: "shadow"`; он не меняет Hermes provider, model или reasoning effort. + +Когда host закончил реальный model call, он записывает outcome через `jev_record_execution`: + +```json +{ + "correlation_id": "из jev_model_route", + "capability_id": "recommended model id", + "status": "completed", + "result": { + "model_route": { + "actual_model_id": "фактически использованный id", + "retry_count": 0, + "outcome": "verified" + } + } +} +``` + +Отчёт только читает receipts: + +```bash +npm run model-route:report -- ~/.hermes/jev-layer/replay/cases.jsonl +``` + +Пока не накоплены репрезентативные outcomes/retries/latency/cost, **не включать** auto-switching моделей. + +### `jev_browser_step` + +Принимает только snapshot, который сделал host: `url`, visible text, targets, tabs, scroll и progress. Возвращает один bounded action: scroll, switch tab, visible link navigation, click, native select или handoff. + +- `scroll` и `switch_tab` могут быть исполнены только через native Hermes browser executor. +- `click`, `select`, `navigate` требуют host approval. +- `submit`, login, payment, delete, секретный ввод и `TYPE_TEXT` не являются быстрым Jev execution path. +- Progress не даёт повторять уже выполненный target и блокирует повторный scroll без измеримого состояния страницы. +- После каждого host шага можно записать routing case и execution receipt; это позволяет replay без повторного browser execution. + +`browser-use/jev-ultrafast` был использован как архитектурный reference (динамическая индексированная action space), но его cloud/browser worker **не установлен и не запущен**. Не подменяй Hermes browser security model его executor'ом. + +### `jev_shadow_compaction` + +Возвращает консервативный report keep/drop для переданного context. Не меняет prompt, не удаляет сообщения и не заменяет host compaction. Pinned requirements, paths, errors и commands сохраняются без provider review; provider error удерживает всё. + +### `jev_supervise` + +Делает bounded judgement о work state. Hermes, не Jev, преобразует его в `continue`, `verify`, `retry`, `finish` или `escalate`. Contradictory evidence не может закончиться `finish`. + +## Secrets и providers + +`integrations/hermes/__init__.py` получает `OPENROUTER_API_KEY` только через Hermes profile-scoped secret store (`agent.secret_scope.get_secret`). Ключ передаётся лишь в короткоживущий локальный Node adapter; не возвращается в tool output и не логируется. + +Для локальных/offline tests используй `provider: "demo"`. Provider-backed проверки — отдельные, opt-in; ключи никогда не клади в repo, plugin manifest, test fixture или обычный shell history. + +## Безопасный change workflow + +1. Прочитай этот документ и `docs/AGENT-IMPLEMENTATION.md`. +2. Проверь target surface: source checkout, installed plugin copy, active profile или global config. +3. Измени source в `/home/hermes/jev-layer` и сначала добавь/измени offline test. +4. Запусти узкий test, затем полный suite. +5. Синхронизируй **только нужные** source files в `~/.hermes/plugins/jev-layer`. +6. Запусти `hermes plugins doctor jev-layer`. +7. Если менялась регистрация/handler/schema/manifest, попроси пользователя о `/restart`. +8. После restart проверь живой gateway tool path и read back exact receipt/report. +9. Не делай `git push`, npm publish, GitHub release или изменение глобального Hermes config без отдельной команды пользователя. + +## Проверки + +Из source checkout: + +```bash +npm test +npm run receipt:mcp-smoke +npm run browser:e2e +npm run model-route:report -- /tmp/nonexistent-cases.jsonl +python3 -m py_compile integrations/hermes/__init__.py integrations/hermes/schemas.py +git diff --check +hermes plugins doctor jev-layer +``` + +Текущее доказанное состояние: `npm test` — 35/35; MCP receipt smoke и browser fixture E2E проходят; gateway-live model route был записан и успешно связан с execution receipt/report. Это не доказательство автоматического роутинга моделей и не утверждение о real-browser speedup. + +## Готовый prompt локальному агенту + +```text +Работай только с jev-layer в указанном target surface. Сначала прочитай: +- docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md +- docs/AGENT-IMPLEMENTATION.md + +Сохрани host ownership: Jev даёт bounded advisory decision; Hermes владеет registry, +permissions, approval, execution, retry/recovery и финальным результатом. Не добавляй +автоматическое переключение моделей, browser executor, глобальный config write, provider +credential или network publication без отдельного явного требования. + +Перед изменением назови source files, installed plugin files и gateway impact. Используй TDD: +сначала узкий offline test, затем implementation, потом npm test + релевантные smoke tests. +Если touch'нуты registration/handler/schema/manifest, синхронизируй plugin copy, прогоняй +hermes plugins doctor jev-layer и попроси /restart. После рестарта сделай live tool call, +запиши/прочитай exact receipt и только затем заявляй, что integration активна. + +Не делай git push, npm publish, release или удаление replay/artifact без отдельной команды. +``` diff --git a/docs/SCHEMA-VERSIONING.md b/docs/SCHEMA-VERSIONING.md index 37c0cfe..31ec475 100644 --- a/docs/SCHEMA-VERSIONING.md +++ b/docs/SCHEMA-VERSIONING.md @@ -21,6 +21,8 @@ The MCP server currently exposes these stable tool names: - `jev_route` - `jev_browser_step` (experimental, opt-in) - `jev_supervise` (experimental, opt-in) +- `jev_model_route` (report-only; experimental, opt-in) +- `jev_shadow_compaction` (report-only; experimental, opt-in) - `jev_record_execution` Their input and structured output are v1. Additive optional properties are compatible. Renaming a tool, changing a required property, changing the meaning of a status, or changing who owns execution requires a new tool/schema version and adapter migration. Keep `tools/list`, `initialize`, and stdio JSON-RPC behavior backward compatible for v1 clients. @@ -66,7 +68,7 @@ A v1 replay file is JSONL. Supported records include: - `record_type: "execution_receipt"` - `record_type: "supervision_case"` -Routing cases retain a sanitized request and a decision summary. Replay reads cases without calling host tools. New optional record fields are compatible; changing record type, correlation semantics, or the meaning of a recorded status requires a new replay schema and migration. +Routing cases retain a sanitized request and a decision summary. `supervision_case.supervision.evidence_state` is an optional v1 object with `state` (`present`, `missing`, or `contradictory`) and bounded source labels. Replay reads cases without calling host tools. New optional record fields are compatible; changing record type, correlation semantics, or the meaning of a recorded status requires a new replay schema and migration. ## Adapter contract diff --git a/integrations/hermes/__init__.py b/integrations/hermes/__init__.py index db0c10c..cb20d4b 100644 --- a/integrations/hermes/__init__.py +++ b/integrations/hermes/__init__.py @@ -16,7 +16,7 @@ from pathlib import Path from typing import Any -from .schemas import BROWSER_STEP, RECORD_EXECUTION, ROUTE, SUPERVISE +from .schemas import BROWSER_STEP, MODEL_ROUTE, RECORD_EXECUTION, ROUTE, SHADOW_COMPACTION, SUPERVISE ROOT = Path(os.environ.get("JEV_LAYER_ROOT", Path(__file__).resolve().parents[2])).expanduser().resolve() CORE = ROOT / "src" / "hermes-adapter.mjs" @@ -38,6 +38,17 @@ def _fallback(reason: str) -> str: def _runtime_env() -> dict[str, str]: env = os.environ.copy() + # Resolve only through Hermes's profile-scoped store, then pass it only to + # this short-lived local Node adapter. It is never returned or logged. + try: + from agent.secret_scope import get_secret + openrouter_key = get_secret("OPENROUTER_API_KEY") + if openrouter_key: + env["OPENROUTER_API_KEY"] = openrouter_key + except Exception: + # The core remains fail-open when the optional provider credential is unavailable. + pass + # Keep replay evidence in the active Hermes profile, never in the source checkout. env.setdefault("JEV_REPLAY_CASES", str(Path(env.get("HERMES_HOME", "~/.hermes")).expanduser() / "jev-layer" / "replay" / "cases.jsonl")) return env @@ -101,6 +112,16 @@ def jev_browser_step(args: dict, **kwargs) -> str: return _route(args, operation="browser_step") +def jev_model_route(args: dict, **kwargs) -> str: + """Recommend a declared model profile; this never changes Hermes's model route.""" + return _route(args, operation="model_route") + + +def jev_shadow_compaction(args: dict, **kwargs) -> str: + """Report Jev's conservative keep/drop candidates without mutating host context.""" + return json.dumps(_call_core("shadow_compaction", dict(args))) + + def jev_supervise(args: dict, **kwargs) -> str: """Return a bounded work-state judgment; Hermes maps it to its own next action.""" request = dict(args) @@ -122,6 +143,8 @@ def jev_record_execution(args: dict, **kwargs) -> str: def register(ctx): ctx.register_tool(name="jev_route", toolset="jev_layer", schema=ROUTE, handler=jev_route) + ctx.register_tool(name="jev_model_route", toolset="jev_layer", schema=MODEL_ROUTE, handler=jev_model_route) + ctx.register_tool(name="jev_shadow_compaction", toolset="jev_layer", schema=SHADOW_COMPACTION, handler=jev_shadow_compaction) ctx.register_tool(name="jev_record_execution", toolset="jev_layer", schema=RECORD_EXECUTION, handler=jev_record_execution) ctx.register_tool(name="jev_supervise", toolset="jev_layer", schema=SUPERVISE, handler=jev_supervise) ctx.register_tool(name="jev_browser_step", toolset="jev_layer", schema=BROWSER_STEP, handler=jev_browser_step) diff --git a/integrations/hermes/plugin.yaml b/integrations/hermes/plugin.yaml index babb00e..8f87fd3 100644 --- a/integrations/hermes/plugin.yaml +++ b/integrations/hermes/plugin.yaml @@ -1,8 +1,10 @@ name: jev-layer version: 0.1.0 -description: Full host-owned Jev routing, receipts, supervision, and browser decisions +description: Full host-owned Jev routing, receipts, supervision, shadow compaction, model routing, and browser decisions provides_tools: - jev_route + - jev_model_route + - jev_shadow_compaction - jev_record_execution - jev_supervise - jev_browser_step diff --git a/integrations/hermes/schemas.py b/integrations/hermes/schemas.py index 851920d..ea73c10 100644 --- a/integrations/hermes/schemas.py +++ b/integrations/hermes/schemas.py @@ -31,6 +31,32 @@ }, } +MODEL_ROUTE = { + "type": "object", + "required": ["intent", "models"], + "properties": { + "intent": {"type": "string"}, + "context": {"type": "object"}, + "models": {"type": "array", "minItems": 1, "items": {"type": "object"}}, + "harness": {"type": "string"}, + "policy": {"type": "object"}, + "provider": {"type": "string", "enum": ["demo", "typesafe", "openrouter"]}, + }, +} + +SHADOW_COMPACTION = { + "type": "object", + "required": ["intent", "context"], + "properties": { + "intent": {"type": "string"}, + "context": {"type": "object"}, + "provider": {"type": "string", "enum": ["demo", "typesafe", "openrouter"]}, + "batch_size": {"type": "integer", "minimum": 1, "maximum": 8}, + "keep_threshold": {"type": "number", "minimum": 0, "maximum": 1}, + "min_confidence": {"type": "number", "minimum": 0, "maximum": 1}, + }, +} + SUPERVISE = { "type": "object", "required": ["job", "observation"], diff --git a/package.json b/package.json index 1c6a72a..ffc17af 100644 --- a/package.json +++ b/package.json @@ -48,6 +48,7 @@ "codex:mcp-smoke": "node scripts/codex-mcp-smoke.mjs", "replay:evaluate": "node scripts/replay-eval.mjs", "receipt:mcp-smoke": "node scripts/mcp-receipt-smoke.mjs", + "model-route:report": "node scripts/model-route-report.mjs", "browser:e2e": "node scripts/browser-e2e.mjs", "browser:benchmark": "node scripts/browser-benchmark.mjs", "supervision:e2e": "node scripts/supervision-e2e.mjs", diff --git a/plugin.yaml b/plugin.yaml index babb00e..8f87fd3 100644 --- a/plugin.yaml +++ b/plugin.yaml @@ -1,8 +1,10 @@ name: jev-layer version: 0.1.0 -description: Full host-owned Jev routing, receipts, supervision, and browser decisions +description: Full host-owned Jev routing, receipts, supervision, shadow compaction, model routing, and browser decisions provides_tools: - jev_route + - jev_model_route + - jev_shadow_compaction - jev_record_execution - jev_supervise - jev_browser_step diff --git a/scripts/mcp-receipt-smoke.mjs b/scripts/mcp-receipt-smoke.mjs index 1b31e03..99b4919 100644 --- a/scripts/mcp-receipt-smoke.mjs +++ b/scripts/mcp-receipt-smoke.mjs @@ -30,6 +30,8 @@ try { await notification("notifications/initialized", {}); const listed = await request(2, "tools/list", {}); assert.ok(listed.result.tools.some((tool) => tool.name === "jev_route")); + assert.ok(listed.result.tools.some((tool) => tool.name === "jev_model_route")); + assert.ok(listed.result.tools.some((tool) => tool.name === "jev_shadow_compaction")); assert.ok(listed.result.tools.some((tool) => tool.name === "jev_record_execution")); const routed = await request(3, "tools/call", { @@ -70,6 +72,19 @@ try { assert.equal(receipt.host.exit_status, 0); assert.equal(receipt.host.duration_ms, 3.2); + const shadow = await request(5, "tools/call", { + name: "jev_shadow_compaction", + arguments: { + provider: "demo", + intent: "inspect context before compaction", + context: { messages: ["stale note", "Requirement: preserve /workspace/config.json"] }, + }, + }); + const shadowReport = shadow.result.structuredContent; + assert.equal(shadowReport.mode, "jev_shadow"); + assert.equal(shadowReport.changed, false); + assert.equal(shadowReport.protected_count, 1); + const records = (await readFile(casesPath, "utf8")).trim().split(/\r?\n/).map(JSON.parse); assert.deepEqual(records.map((record) => record.record_type), ["routing_case", "execution_receipt"]); console.log(JSON.stringify({ ok: true, cases_path: casesPath, receipt }, null, 2)); diff --git a/scripts/model-route-report.mjs b/scripts/model-route-report.mjs new file mode 100644 index 0000000..e800846 --- /dev/null +++ b/scripts/model-route-report.mjs @@ -0,0 +1,15 @@ +#!/usr/bin/env node +import { readFile } from "node:fs/promises"; +import { buildModelRouteReport } from "../src/model-route-metrics.mjs"; + +const path = process.argv[2] ?? process.env.JEV_REPLAY_CASES ?? ".jev/replay/cases.jsonl"; +let text = ""; +try { + text = await readFile(path, "utf8"); +} catch (error) { + if (error?.code !== "ENOENT") throw error; +} +const records = text.split(/\r?\n/).filter(Boolean).flatMap((line) => { + try { return [JSON.parse(line)]; } catch { return []; } +}); +console.log(JSON.stringify(buildModelRouteReport(records), null, 2)); diff --git a/scripts/supervision-e2e.mjs b/scripts/supervision-e2e.mjs index 1eb0029..39fc63a 100644 --- a/scripts/supervision-e2e.mjs +++ b/scripts/supervision-e2e.mjs @@ -66,9 +66,38 @@ try { assert.equal(result.metrics.jev_calls, 1); assert.ok(result.receipt.correlation_id); + const contradictory = await request(5, "tools/call", { + name: "jev_supervise", + arguments: { + enabled: true, + provider: process.env.JEV_SUPERVISION_PROVIDER ?? "demo", + harness: "supervision-e2e", + job: { requirements: ["run tests"] }, + observation: { status: "claimed-complete" }, + evidence: { + tests_passed: true, + tests_failed: true, + judgments: { + requirements_addressed: 0.9, + verification_needed: 0.1, + meaningful_progress: 0.9, + worker_stuck: 0.05, + work_off_track: 0.05, + completion: 0.9, + }, + }, + }, + }); + const contradictoryResult = contradictory.result.structuredContent; + assert.equal(contradictoryResult.status, "judged"); + assert.equal(contradictoryResult.action, "verify"); + assert.equal(contradictoryResult.evidence_state.state, "contradictory"); + const records = (await readFile(casesPath, "utf8")).trim().split(/\r?\n/).map(JSON.parse); - assert.deepEqual(records.map((record) => record.record_type), ["supervision_case"]); + assert.deepEqual(records.map((record) => record.record_type), ["supervision_case", "supervision_case"]); assert.equal(records[0].supervision.action, "finish"); + assert.equal(records[1].supervision.action, "verify"); + assert.equal(records[1].supervision.evidence_state.state, "contradictory"); const replayed = deterministicSupervisionPolicy({ assessment: records[0].supervision.assessment, evidence: records[0].request.context.evidence, diff --git a/src/hermes-adapter.mjs b/src/hermes-adapter.mjs index 3050047..90f929f 100644 --- a/src/hermes-adapter.mjs +++ b/src/hermes-adapter.mjs @@ -12,6 +12,8 @@ import { configuredProvider, configuredReplayPath, loadConfig } from "./config.m import { appendExecutionReceipt, appendRoutingCase, buildExecutionReceipt, replayCasePath } from "./receipts.mjs"; import { routeRequest } from "./route.mjs"; import { superviseWork } from "./supervision.mjs"; +import { buildShadowCompactionReport } from "./shadow-compaction.mjs"; +import { recommendModelRoute } from "./model-routing.mjs"; const { config } = await loadConfig(); const casesPath = replayCasePath(configuredReplayPath(config)); @@ -62,6 +64,30 @@ async function handle({ operation, args = {}, decision = null } = {}) { receiptPath: casesPath, }); } + if (operation === "model_route") { + const { provider: _provider, ...input } = args; + const result = await recommendModelRoute({ ...input, provider }); + await persistRoutingCase({ + schema_version: 1, + harness: input.harness ?? "hermes", + intent: input.intent, + context: { ...(input.context ?? {}), model_route: { mode: "shadow" } }, + capabilities: (input.models ?? []).map((profile) => ({ + id: profile.id, + kind: "model", + name: profile.model ?? profile.id, + description: profile.description ?? "", + source: profile.provider ?? null, + risk: "low", + })), + policy: input.policy ?? {}, + }, result); + return result; + } + if (operation === "shadow_compaction") { + const { provider: _provider, ...input } = args; + return buildShadowCompactionReport({ ...input, provider }); + } if (operation === "record_execution") { if (!decision || typeof decision !== "object") throw new TypeError("decision is required for record_execution"); if (args.correlation_id !== decision.correlation_id) throw new TypeError("correlation_id does not match the routed decision"); diff --git a/src/mcp-server.mjs b/src/mcp-server.mjs index e311528..92c1be6 100755 --- a/src/mcp-server.mjs +++ b/src/mcp-server.mjs @@ -4,6 +4,8 @@ import { appendExecutionReceipt, appendRoutingCase, buildExecutionReceipt, repla import { configuredProvider, configuredReplayPath, loadConfig } from "./config.mjs"; import { decideBrowserStep } from "./browser.mjs"; import { superviseWork } from "./supervision.mjs"; +import { buildShadowCompactionReport } from "./shadow-compaction.mjs"; +import { recommendModelRoute } from "./model-routing.mjs"; import { routeRequest } from "./route.mjs"; const { config } = await loadConfig(); @@ -65,6 +67,38 @@ const tools = [ }, }, }, + { + name: "jev_model_route", + description: "Recommend one host-declared model profile for the next call. This is shadow-only: it never changes provider, model, reasoning, or execution.", + inputSchema: { + type: "object", + required: ["intent", "models"], + properties: { + intent: { type: "string" }, + context: { type: "object" }, + models: { type: "array", minItems: 1, items: { type: "object" } }, + harness: { type: "string" }, + policy: { type: "object" }, + provider: { type: "string", enum: ["demo", "typesafe", "openrouter"] }, + }, + }, + }, + { + name: "jev_shadow_compaction", + description: "Report conservative Jev keep/drop candidates for host-supplied context. This tool never changes, summarizes, or deletes context.", + inputSchema: { + type: "object", + required: ["intent", "context"], + properties: { + intent: { type: "string" }, + context: { type: "object" }, + provider: { type: "string", enum: ["demo", "typesafe", "openrouter"] }, + batch_size: { type: "integer", minimum: 1, maximum: 8 }, + keep_threshold: { type: "number", minimum: 0, maximum: 1 }, + min_confidence: { type: "number", minimum: 0, maximum: 1 }, + }, + }, + }, { name: "jev_record_execution", description: "Attach a host execution result to a Jev decision and persist the unified execution receipt.", @@ -115,6 +149,8 @@ async function handle(message) { if (params.name === "jev_route") return route(args, id); if (params.name === "jev_browser_step") return browserStep(args, id); if (params.name === "jev_supervise") return supervise(args, id); + if (params.name === "jev_model_route") return modelRoute(args, id); + if (params.name === "jev_shadow_compaction") return shadowCompaction(args, id); if (params.name === "jev_record_execution") return recordExecution(args, id); return jsonRpcError(id, -32602, `unknown tool: ${params.name}`); } @@ -162,6 +198,24 @@ async function supervise(args, id) { return toolResult(id, result, false); } +async function modelRoute(args, id) { + const { provider: requestedProvider, ...input } = args; + const result = await recommendModelRoute({ + ...input, + provider: configuredProvider(config, requestedProvider), + }); + return toolResult(id, result, false); +} + +async function shadowCompaction(args, id) { + const { provider: requestedProvider, ...input } = args; + const result = await buildShadowCompactionReport({ + ...input, + provider: configuredProvider(config, requestedProvider), + }); + return toolResult(id, result, false); +} + async function persistDecision(request, decision, id) { pendingDecisions.set(decision.correlation_id, { request, decision }); try { diff --git a/src/model-route-metrics.mjs b/src/model-route-metrics.mjs new file mode 100644 index 0000000..50bf75a --- /dev/null +++ b/src/model-route-metrics.mjs @@ -0,0 +1,89 @@ +/** + * Read-only analysis for model-route shadow cases. Recommendations stay + * advisory: this module joins evidence; it never changes host model settings. + */ +export function buildModelRouteReport(records = []) { + const routes = new Map(); + const executions = new Map(); + for (const record of Array.isArray(records) ? records : []) { + if (!record || typeof record !== "object" || typeof record.correlation_id !== "string") continue; + if (record.record_type === "routing_case" && record.request?.context?.model_route?.mode === "shadow") routes.set(record.correlation_id, record); + if (record.record_type === "execution_receipt") executions.set(record.correlation_id, record); + } + + const rows = [...routes.values()].map((route) => toRow(route, executions.get(route.correlation_id))).sort((a, b) => a.recommended_model_id.localeCompare(b.recommended_model_id)); + return { + schema_version: 1, + mode: "shadow", + generated_at: new Date().toISOString(), + summary: summarize(rows), + models: groupByRecommendation(rows), + }; +} + +function toRow(route, execution) { + const result = execution?.host?.result?.model_route; + const recommended = string(route.decision?.selected, "unrecommended"); + const actual = string(result?.actual_model_id, null); + const outcome = string(result?.outcome, execution ? string(execution.host?.status, "unknown") : null); + return { + correlation_id: route.correlation_id, + recommended_model_id: recommended, + actual_model_id: actual, + recommendation_match: actual !== null && actual === recommended, + outcome, + retry_count: boundedInteger(result?.retry_count), + host_duration_ms: finite(execution?.host?.duration_ms), + jev_latency_ms: finite(route.decision?.receipt?.latency_ms), + jev_cost_usd: finite(route.decision?.receipt?.cost_usd), + }; +} + +function summarize(rows) { + const matched = rows.filter((row) => row.recommendation_match).length; + const actual = rows.filter((row) => row.actual_model_id !== null).length; + return { + decisions: rows.length, + with_host_outcome: rows.filter((row) => row.outcome !== null).length, + with_actual_model: actual, + recommendation_matches: matched, + recommendation_match_rate: actual ? matched / actual : null, + verified_outcomes: rows.filter((row) => row.outcome === "verified" || row.outcome === "completed").length, + failed_outcomes: rows.filter((row) => row.outcome === "failed").length, + total_retry_count: rows.reduce((sum, row) => sum + (row.retry_count ?? 0), 0), + total_jev_cost_usd: sum(rows.map((row) => row.jev_cost_usd)), + average_jev_latency_ms: average(rows.map((row) => row.jev_latency_ms)), + average_host_duration_ms: average(rows.map((row) => row.host_duration_ms)), + }; +} + +function groupByRecommendation(rows) { + const grouped = new Map(); + for (const row of rows) { + const group = grouped.get(row.recommended_model_id) ?? { + recommended_model_id: row.recommended_model_id, + decisions: 0, + with_host_outcome: 0, + recommendation_matches: 0, + verified_outcomes: 0, + failed_outcomes: 0, + total_retry_count: 0, + host_actual_model_ids: {}, + }; + group.decisions += 1; + if (row.outcome !== null) group.with_host_outcome += 1; + if (row.recommendation_match) group.recommendation_matches += 1; + if (row.outcome === "verified" || row.outcome === "completed") group.verified_outcomes += 1; + if (row.outcome === "failed") group.failed_outcomes += 1; + group.total_retry_count += row.retry_count ?? 0; + if (row.actual_model_id) group.host_actual_model_ids[row.actual_model_id] = (group.host_actual_model_ids[row.actual_model_id] ?? 0) + 1; + grouped.set(row.recommended_model_id, group); + } + return [...grouped.values()].sort((a, b) => a.recommended_model_id.localeCompare(b.recommended_model_id)); +} + +function string(value, fallback) { return typeof value === "string" && value.trim() ? value.trim() : fallback; } +function finite(value) { return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null; } +function boundedInteger(value) { return Number.isInteger(value) && value >= 0 ? value : null; } +function sum(values) { return values.reduce((total, value) => total + (value ?? 0), 0); } +function average(values) { const present = values.filter((value) => value !== null); return present.length ? sum(present) / present.length : null; } diff --git a/src/model-routing.mjs b/src/model-routing.mjs new file mode 100644 index 0000000..fdca461 --- /dev/null +++ b/src/model-routing.mjs @@ -0,0 +1,86 @@ +import { routeRequest } from "./route.mjs"; + +/** + * Recommend one host-declared model profile for the next model call. + * This is a shadow decision: it never mutates the host model/provider setting. + */ +export async function recommendModelRoute({ + intent, + context = {}, + models, + harness = "unknown", + policy = {}, + provider = "demo", +} = {}) { + if (!Array.isArray(models) || models.length === 0) throw new TypeError("models[] is required"); + const profiles = models.map(normalizeProfile); + const profileById = new Map(profiles.map((profile) => [profile.id, profile])); + const decision = await routeRequest({ + harness, + intent, + context: { ...context, model_route: { mode: "shadow", profiles: profiles.map(publicProfile) } }, + capabilities: profiles.map(toCapability), + policy: { + ...policy, + // A recommendation must never create an approval boundary by itself. + confirmation_risk_levels: [], + }, + }, { provider }); + return { + ...decision, + route_mode: "shadow", + recommended_model: decision.selected ? publicProfile(profileById.get(decision.selected)) : null, + }; +} + +function normalizeProfile(value, index) { + if (!value || typeof value !== "object") throw new TypeError(`models[${index}] must be an object`); + const id = required(value.id, `models[${index}].id`); + const model = required(value.model ?? value.id, `models[${index}].model`); + const provider = string(value.provider, "unknown"); + const reasoning_effort = string(value.reasoning_effort, "default"); + const description = string(value.description, `${provider}/${model}; reasoning=${reasoning_effort}`); + return { + id, + provider, + model, + reasoning_effort, + description, + available: value.available !== false && value.availability?.available !== false, + availability_reason: value.availability?.reason ?? null, + }; +} + +function toCapability(profile) { + return { + id: profile.id, + kind: "model", + name: profile.model, + description: profile.description, + risk: "low", + available: profile.available, + availability: { available: profile.available, reason: profile.availability_reason }, + source: profile.provider, + metadata: { model: profile.model }, + }; +} + +function publicProfile(profile) { + if (!profile) return null; + return { + id: profile.id, + provider: profile.provider, + model: profile.model, + reasoning_effort: profile.reasoning_effort, + available: profile.available, + }; +} + +function required(value, label) { + if (typeof value !== "string" || !value.trim()) throw new TypeError(`${label} is required`); + return value.trim(); +} + +function string(value, fallback) { + return typeof value === "string" && value.trim() ? value.trim() : fallback; +} diff --git a/src/receipts.mjs b/src/receipts.mjs index 6a69d17..1693dea 100644 --- a/src/receipts.mjs +++ b/src/receipts.mjs @@ -101,6 +101,7 @@ export function buildSupervisionReceipt({ request, result, recordedAt = new Date action: result.action ?? "continue", reason: result.reason ?? null, assessment: result.assessment ?? null, + evidence_state: result.evidence_state ?? request.context?.supervision?.evidence_state ?? null, policy: result.policy ?? null, jev: { provider: result.receipt.provider ?? null, diff --git a/src/shadow-compaction.mjs b/src/shadow-compaction.mjs new file mode 100644 index 0000000..6aaa8a1 --- /dev/null +++ b/src/shadow-compaction.mjs @@ -0,0 +1,178 @@ +import { sanitizeContext } from "./context-filter.mjs"; +import { isPinnedEvidence } from "./relevance-filter.mjs"; +import { OpenRouterDecisionsProvider, TypeSafeProvider } from "./providers/typesafe.mjs"; +import { DemoProvider } from "./providers/demo.mjs"; + +const LIST_KEYS = new Set(["messages", "events", "logs", "tool_results", "history", "transcript"]); +const DEFAULT_BATCH_SIZE = 8; +const DEFAULT_KEEP_THRESHOLD = 0.5; +const DEFAULT_MIN_CONFIDENCE = 0.7; +const MAX_ITEMS = 64; +const MAX_PREVIEW = 320; + +/** + * Produce advisory Jev compaction decisions without mutating context. + * Pinned evidence is never sent as droppable and provider failure retains all items. + */ +export async function buildShadowCompactionReport({ + intent = "Assess which context records must remain available to a later agent.", + context = {}, + provider = "demo", + batchSize = DEFAULT_BATCH_SIZE, + keepThreshold = DEFAULT_KEEP_THRESHOLD, + minConfidence = DEFAULT_MIN_CONFIDENCE, +} = {}) { + const items = collectContextItems(context); + const normalizedBatchSize = boundedInteger(batchSize, DEFAULT_BATCH_SIZE, 1, DEFAULT_BATCH_SIZE); + const normalizedThreshold = boundedNumber(keepThreshold, DEFAULT_KEEP_THRESHOLD); + const normalizedConfidence = boundedNumber(minConfidence, DEFAULT_MIN_CONFIDENCE); + const protectedItems = items.filter((item) => item.pinned); + const evaluable = items.filter((item) => !item.pinned); + const decisions = new Map(protectedItems.map((item) => [item.id, { + decision: "keep", + probability: 1, + confidence: 1, + reason: `pinned:${item.pinned_reasons.join(",")}`, + }])); + let costUsd = 0; + let failed = null; + + try { + const resolved = resolveProvider(provider); + if (!resolved || typeof resolved.evaluate !== "function") throw new Error("provider does not support batched evaluation"); + for (const batch of batches(evaluable, normalizedBatchSize)) { + const raw = await resolved.evaluate({ + state: { + intent: boundedText(intent), + compaction: { + mode: "shadow", + instruction: "For each item, decide whether dropping it would lose a requirement, exact identifier, path, command, error, user decision, execution receipt, or other evidence needed later. Be conservative: keep when uncertain.", + items: batch.map(publicItem), + }, + }, + questions: Object.fromEntries(batch.map((item, index) => [`item_${index}`, keepQuestion()])), + }); + costUsd += finiteCost(raw?.usage?.cost) ?? 0; + for (const [index, item] of batch.entries()) { + const answer = readNoul(raw?.answers?.[`item_${index}`]); + if (!answer) throw new Error(`provider response has no valid item_${index} judgment`); + const decision = answer.confidence < normalizedConfidence || answer.probability >= normalizedThreshold ? "keep" : "drop_candidate"; + decisions.set(item.id, { + decision, + probability: answer.probability, + confidence: answer.confidence, + reason: answer.confidence < normalizedConfidence ? "low_confidence" : decision === "keep" ? "provider_keep" : "provider_drop_candidate", + }); + } + } + } catch (error) { + failed = error instanceof Error ? error.message : String(error); + for (const item of evaluable) { + decisions.set(item.id, { decision: "keep", probability: null, confidence: null, reason: "provider_error" }); + } + } + + const reportedItems = items.map((item) => ({ ...publicItem(item), ...decisions.get(item.id) })); + const droppableCount = reportedItems.filter((item) => item.decision === "drop_candidate").length; + return { + schema_version: 1, + mode: "jev_shadow", + status: failed ? "fallback" : "judged", + changed: false, + candidate_count: reportedItems.length, + evaluated_count: evaluable.length, + protected_count: protectedItems.length, + retained_count: reportedItems.length - droppableCount, + droppable_count: droppableCount, + keep_threshold: normalizedThreshold, + min_confidence: normalizedConfidence, + cost_usd: failed ? null : Number(costUsd.toFixed(8)), + fallback: failed ? { type: "provider_error", reason: failed } : null, + items: reportedItems, + }; +} + +export function collectContextItems(context = {}) { + const source = context && typeof context === "object" ? context : {}; + const items = []; + for (const [list, value] of Object.entries(source)) { + if (!LIST_KEYS.has(list) || !Array.isArray(value)) continue; + for (const [index, original] of value.entries()) { + if (items.length >= MAX_ITEMS) return items; + const sanitized = sanitizeContext(original); + const evidence = isPinnedEvidence(sanitized); + items.push({ + id: `${list}[${index}]`, + list, + index, + value: sanitized, + pinned: evidence.pinned, + pinned_reasons: evidence.reasons, + }); + } + } + return items; +} + +function publicItem(item) { + return { + id: item.id, + list: item.list, + index: item.index, + pinned: item.pinned, + pinned_reasons: item.pinned_reasons, + preview: boundedText(typeof item.value === "string" ? item.value : JSON.stringify(item.value)), + }; +} + +function keepQuestion() { + return { + type: "noul", + instructions: "Would dropping this context record lose information needed by a later agent? Return true to keep it and false only when it is safely stale or redundant.", + criteria: { + true: "Keep: it contains a requirement, decision, identifier, path, command, error, execution evidence, or context that may matter later; uncertainty means keep.", + false: "Drop candidate: it is safely stale or redundant and losing it would not affect later work.", + }, + }; +} + +function readNoul(answer) { + if (!answer || typeof answer !== "object") return null; + const probability = ["probability", "p_true", "score", "value", "noul"].map((key) => answer[key]).find(validProbability); + if (!validProbability(probability)) return null; + const confidence = validProbability(answer.confidence) ? answer.confidence : probability; + return { probability, confidence }; +} + +function resolveProvider(provider) { + if (provider && typeof provider === "object") return provider; + if (provider === "demo") return new DemoProvider(); + if (provider === "openrouter") return new OpenRouterDecisionsProvider(); + if (provider === "typesafe") return new TypeSafeProvider(); + throw new Error(`unsupported provider: ${provider}`); +} + +function batches(items, size) { + return Array.from({ length: Math.ceil(items.length / size) }, (_, index) => items.slice(index * size, (index + 1) * size)); +} + +function boundedText(value) { + const text = String(value ?? ""); + return text.length > MAX_PREVIEW ? `${text.slice(0, MAX_PREVIEW)}…` : text; +} + +function boundedNumber(value, fallback) { + return typeof value === "number" && Number.isFinite(value) ? Math.max(0, Math.min(1, value)) : fallback; +} + +function boundedInteger(value, fallback, min, max) { + return Number.isInteger(value) ? Math.max(min, Math.min(max, value)) : fallback; +} + +function validProbability(value) { + return typeof value === "number" && Number.isFinite(value) && value >= 0 && value <= 1; +} + +function finiteCost(value) { + return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null; +} diff --git a/src/supervision.mjs b/src/supervision.mjs index 561ba21..d3347c1 100644 --- a/src/supervision.mjs +++ b/src/supervision.mjs @@ -41,6 +41,7 @@ export function buildSupervisionRequest({ evidence: bounded(evidence), supervision: { judgments: bounded(evidence?.judgments ?? {}), + evidence_state: deterministicEvidenceState(evidence), }, }, }; @@ -70,12 +71,42 @@ export function normalizeAssessment(raw) { return assessment; } +export function deterministicEvidenceState(evidence = {}) { + const sources = []; + const positive = evidence?.verification_passed === true || evidence?.tests_passed === true; + const negative = evidence?.verification_failed === true || evidence?.tests_failed === true; + if (evidence?.verification_passed === true) sources.push("verification_passed"); + if (evidence?.tests_passed === true) sources.push("tests_passed"); + if (evidence?.verification_failed === true) sources.push("verification_failed"); + if (evidence?.tests_failed === true) sources.push("tests_failed"); + if (Array.isArray(evidence?.receipts)) { + for (const receipt of evidence.receipts) { + if (!receipt || typeof receipt !== "object") continue; + const id = typeof receipt.id === "string" && receipt.id ? receipt.id : null; + if (receipt.status === "completed") { + sources.push(id ? `receipt:${id}` : "receipt:completed"); + } else if (receipt.status === "failed") { + sources.push(id ? `receipt_failed:${id}` : "receipt:failed"); + } + } + } + const hasPositiveReceipt = sources.some((source) => source === "receipt:completed" || source.startsWith("receipt:")); + const hasNegativeReceipt = sources.some((source) => source === "receipt:failed" || source.startsWith("receipt_failed:")); + if ((positive || hasPositiveReceipt) && (negative || hasNegativeReceipt)) return { state: "contradictory", sources }; + if (negative || hasNegativeReceipt) return { state: "contradictory", sources }; + return { state: positive || hasPositiveReceipt ? "present" : "missing", sources }; +} + export function deterministicSupervisionPolicy({ assessment, evidence = {}, attempts = 0, policy = {} } = {}) { const values = assessment && typeof assessment === "object" ? assessment : {}; const threshold = Number.isFinite(policy.threshold) ? Math.max(0, Math.min(1, policy.threshold)) : 0.7; const maxRetries = Number.isInteger(policy.max_retries) && policy.max_retries >= 0 ? policy.max_retries : 2; - const verificationEvidence = evidence.verification_passed === true || evidence.tests_passed === true; + const evidenceState = deterministicEvidenceState(evidence); + const verificationEvidence = evidenceState.state === "present"; + if (evidenceState.state === "contradictory") { + return { action: "verify", reason: "verification evidence is contradictory", policy_source: "deterministic_host_policy", evidence_state: evidenceState }; + } if (value(values.worker_stuck) >= threshold) { return attempts < maxRetries ? { action: "retry", reason: "worker appears stuck and retry budget remains", policy_source: "deterministic_host_policy" } @@ -150,6 +181,7 @@ export async function superviseWork({ action: selectedPolicy.action, reason: selectedPolicy.reason, assessment, + evidence_state: deterministicEvidenceState(evidence), policy: selectedPolicy, request, metrics: { diff --git a/test/hermes-adapter.test.mjs b/test/hermes-adapter.test.mjs index bfce125..9f48c9a 100644 --- a/test/hermes-adapter.test.mjs +++ b/test/hermes-adapter.test.mjs @@ -44,15 +44,19 @@ test("Hermes adapter routes a closed candidate set and records a correlated rece assert.equal(records.filter((record) => record.record_type === "execution_receipt").length, 1); }); -test("Hermes adapter exposes supervision and browser surfaces without execution", async () => { +test("Hermes adapter exposes supervision, shadow-compaction, and browser surfaces without execution", async () => { const dir = await mkdtemp(join(tmpdir(), "jev-hermes-adapter-")); const replayPath = join(dir, "cases.jsonl"); - const [supervision, browser] = await run([ + const [supervision, shadow, browser] = await run([ { operation: "supervise", args: { provider: "demo", enabled: true, harness: "hermes-test", job: { goal: "verify" }, observation: { state: "done" }, evidence: { tests_passed: true } } }, + { operation: "shadow_compaction", args: { provider: "demo", intent: "review context", context: { messages: ["old note", "Requirement: preserve /workspace/project/config.json"] } } }, { operation: "browser_step", args: { provider: "demo", enabled: true, harness: "hermes-test", goal: "Open the visible docs link", observation: { url: "https://example.test", targets: [{ id: "docs", name: "Docs", href: "/docs", visible: true, clickable: true }], tabs: [], scroll: { down: false, up: false } } } }, ], replayPath); assert.equal(supervision.status, "judged"); assert.ok(["continue", "verify", "retry", "finish", "escalate"].includes(supervision.action)); + assert.equal(shadow.mode, "jev_shadow"); + assert.equal(shadow.changed, false); + assert.equal(shadow.protected_count, 1); assert.ok(["selected", "fallback", "no_decision", "needs_confirmation"].includes(browser.status)); assert.equal(browser.execution.enabled, false); }); diff --git a/test/model-route-metrics.test.mjs b/test/model-route-metrics.test.mjs new file mode 100644 index 0000000..173df0e --- /dev/null +++ b/test/model-route-metrics.test.mjs @@ -0,0 +1,61 @@ +import assert from "node:assert/strict"; +import { test } from "node:test"; +import { buildModelRouteReport } from "../src/model-route-metrics.mjs"; + +const rows = [ + { + record_type: "routing_case", + correlation_id: "one", + request: { context: { model_route: { mode: "shadow", profiles: [{ id: "fast" }, { id: "frontier" }] } } }, + decision: { status: "selected", selected: "fast", receipt: { latency_ms: 4, cost_usd: 0.001 } }, + }, + { + record_type: "execution_receipt", + correlation_id: "one", + host: { + status: "completed", + duration_ms: 20, + result: { model_route: { actual_model_id: "fast", retry_count: 1, outcome: "verified" } }, + }, + }, + { + record_type: "routing_case", + correlation_id: "two", + request: { context: { model_route: { mode: "shadow", profiles: [{ id: "fast" }, { id: "frontier" }] } } }, + decision: { status: "selected", selected: "frontier", receipt: { latency_ms: 6, cost_usd: 0.002 } }, + }, + { + record_type: "execution_receipt", + correlation_id: "two", + host: { + status: "failed", + duration_ms: 30, + result: { model_route: { actual_model_id: "fast", retry_count: 2, outcome: "failed" } }, + }, + }, +]; + +test("model route report joins shadow recommendation to host outcome without treating it as a route change", () => { + const report = buildModelRouteReport(rows); + assert.equal(report.schema_version, 1); + assert.equal(report.summary.decisions, 2); + assert.equal(report.summary.with_host_outcome, 2); + assert.equal(report.summary.recommendation_match_rate, 0.5); + assert.equal(report.summary.verified_outcomes, 1); + assert.equal(report.summary.failed_outcomes, 1); + assert.equal(report.summary.total_retry_count, 3); + assert.deepEqual(report.models.map((row) => row.recommended_model_id), ["fast", "frontier"]); + assert.equal(report.models[0].host_actual_model_ids.fast, 1); + assert.equal(report.models[1].host_actual_model_ids.fast, 1); + assert.equal(report.models[1].recommendation_matches, 0); +}); + +test("model route report fails open on incomplete or unrelated receipts", () => { + const report = buildModelRouteReport([ + { record_type: "routing_case", correlation_id: "unrelated", request: { context: {} }, decision: { selected: "tool" } }, + { record_type: "routing_case", correlation_id: "pending", request: { context: { model_route: { mode: "shadow" } } }, decision: { selected: null } }, + ]); + assert.equal(report.summary.decisions, 1); + assert.equal(report.summary.with_host_outcome, 0); + assert.equal(report.models[0].recommended_model_id, "unrecommended"); +}); diff --git a/test/model-routing.test.mjs b/test/model-routing.test.mjs new file mode 100644 index 0000000..eadb7a8 --- /dev/null +++ b/test/model-routing.test.mjs @@ -0,0 +1,59 @@ +import assert from "node:assert/strict"; +import { test } from "node:test"; +import { recommendModelRoute } from "../src/model-routing.mjs"; + +const models = [ + { + id: "fast", + provider: "openrouter", + model: "small-model", + reasoning_effort: "low", + description: "fast and cheap for simple classification or file inspection", + }, + { + id: "frontier", + provider: "openai-codex", + model: "frontier-model", + reasoning_effort: "high", + description: "strong model for difficult implementation and ambiguous reasoning", + }, +]; + +test("model routing returns an advisory profile and never changes the host route", async () => { + let captured; + const result = await recommendModelRoute({ + intent: "classify a small batch of already structured feedback", + context: { task_size: "small", external_side_effects: false }, + models, + provider: { + name: "model-route-test-provider", + async decide(input) { + captured = input; + return { + answers: { tool: { type: "choice", choice: "fast", probabilities: { fast: 0.92, frontier: 0.08 }, confidence: 0.92 } }, + usage: { cost: 0.0001 }, + }; + }, + }, + }); + + assert.equal(result.route_mode, "shadow"); + assert.equal(result.status, "selected"); + assert.equal(result.selected, "fast"); + assert.equal(result.recommended_model.model, "small-model"); + assert.equal(result.execution.enabled, false); + assert.equal(captured.candidates.every((candidate) => candidate.kind === "model"), true); +}); + +test("model routing excludes unavailable profiles and fails open with no selected profile", async () => { + const result = await recommendModelRoute({ + intent: "inspect a file", + models: [{ ...models[0], available: false }], + provider: "demo", + }); + + assert.equal(result.status, "no_decision"); + assert.equal(result.selected, null); + assert.equal(result.recommended_model, null); + assert.equal(result.execution.enabled, false); +}); diff --git a/test/shadow-compaction.test.mjs b/test/shadow-compaction.test.mjs new file mode 100644 index 0000000..1dd17af --- /dev/null +++ b/test/shadow-compaction.test.mjs @@ -0,0 +1,67 @@ +import assert from "node:assert/strict"; +import { test } from "node:test"; +import { buildShadowCompactionReport } from "../src/shadow-compaction.mjs"; + +const context = { + messages: [ + "stale casual discussion", + "Requirement: preserve /workspace/project/config.json exactly", + "recent working note", + ], + tool_results: [ + "plain old tool output", + ], +}; + +test("shadow compaction is report-only, preserves pinned evidence, and exposes candidate decisions", async () => { + const before = JSON.stringify(context); + let received; + const report = await buildShadowCompactionReport({ + intent: "prepare a safe compaction report", + context, + provider: { + name: "shadow-test-provider", + async evaluate(input) { + received = input; + return { + answers: { + item_0: { type: "noul", probability: 0.1, confidence: 0.9 }, + item_1: { type: "noul", probability: 0.9, confidence: 0.9 }, + item_2: { type: "noul", probability: 0.4, confidence: 0.2 }, + }, + usage: { cost: 0.0002 }, + }; + }, + }, + }); + + assert.equal(JSON.stringify(context), before); + assert.equal(report.mode, "jev_shadow"); + assert.equal(report.status, "judged"); + assert.equal(report.changed, false); + assert.equal(report.candidate_count, 4); + assert.equal(report.protected_count, 1); + assert.equal(report.droppable_count, 1); + assert.equal(report.retained_count, 3); + assert.ok(received.questions.item_0); + assert.equal(report.items.find((item) => item.id === "messages[1]").decision, "keep"); + assert.equal(report.items.find((item) => item.id === "messages[1]").reason, "pinned:path,requirement"); + assert.equal(report.items.find((item) => item.id === "messages[0]").decision, "drop_candidate"); + assert.equal(report.items.find((item) => item.id === "tool_results[0]").decision, "keep"); + assert.equal(report.items.find((item) => item.id === "tool_results[0]").reason, "low_confidence"); +}); + +test("shadow compaction fails open and marks all unresolved items as retained", async () => { + const report = await buildShadowCompactionReport({ + intent: "prepare a safe compaction report", + context: { messages: ["old note"] }, + provider: { name: "offline", async evaluate() { throw new Error("provider unavailable"); } }, + }); + + assert.equal(report.status, "fallback"); + assert.equal(report.changed, false); + assert.equal(report.droppable_count, 0); + assert.equal(report.retained_count, 1); + assert.equal(report.items[0].decision, "keep"); + assert.equal(report.items[0].reason, "provider_error"); +}); diff --git a/test/supervision.test.mjs b/test/supervision.test.mjs index 330cb05..16e891d 100644 --- a/test/supervision.test.mjs +++ b/test/supervision.test.mjs @@ -4,6 +4,7 @@ import { tmpdir } from "node:os"; import { join } from "node:path"; import { test } from "node:test"; import { + deterministicEvidenceState, deterministicSupervisionPolicy, normalizeAssessment, superviseWork, @@ -52,11 +53,26 @@ test("supervision contract exposes bounded dimensions and parses Noul answers", }); }); +test("deterministic evidence state distinguishes proof, absence, and contradiction", () => { + assert.deepEqual(deterministicEvidenceState({}), { state: "missing", sources: [] }); + assert.deepEqual(deterministicEvidenceState({ tests_passed: true, receipts: [{ id: "test:1", status: "completed" }] }), { + state: "present", + sources: ["tests_passed", "receipt:test:1"], + }); + assert.deepEqual(deterministicEvidenceState({ tests_passed: true, tests_failed: true }), { + state: "contradictory", + sources: ["tests_passed", "tests_failed"], + }); +}); + test("deterministic host policy owns the supervision action", () => { assert.equal(deterministicSupervisionPolicy({ assessment: assessment({ verification_needed: 0.9 }) }).action, "verify"); assert.equal(deterministicSupervisionPolicy({ assessment: assessment({ worker_stuck: 0.9 }), attempts: 0 }).action, "retry"); assert.equal(deterministicSupervisionPolicy({ assessment: assessment({ worker_stuck: 0.9 }), attempts: 2 }).action, "escalate"); assert.equal(deterministicSupervisionPolicy({ assessment: assessment(), evidence: { tests_passed: true } }).action, "finish"); + const contradiction = deterministicSupervisionPolicy({ assessment: assessment(), evidence: { tests_passed: true, tests_failed: true } }); + assert.equal(contradiction.action, "verify"); + assert.match(contradiction.reason, /contradictory/); }); test("supervision is opt-in, fail-open, and records replayable judgments", async () => { @@ -91,9 +107,11 @@ test("supervision is opt-in, fail-open, and records replayable judgments", async assert.equal(result.action, "finish"); assert.equal(result.metrics.jev_calls, 1); assert.equal(result.metrics.cost_usd, 0.0003); + assert.deepEqual(result.evidence_state, { state: "present", sources: ["tests_passed"] }); const cases = await readSupervisionCases(path); assert.equal(cases.length, 1); assert.equal(cases[0].record_type, "supervision_case"); assert.equal(cases[0].supervision.action, "finish"); + assert.deepEqual(cases[0].supervision.evidence_state, { state: "present", sources: ["tests_passed"] }); assert.equal(JSON.parse(await readFile(path, "utf8")).record_type, "supervision_case"); }); From 801d80ce9526a84e85750a60df650099ed3c091c Mon Sep 17 00:00:00 2001 From: typakon4 Date: Mon, 21 Sep 2026 18:19:14 +0300 Subject: [PATCH 2/3] fix(hermes): preserve configured provider options --- src/browser.mjs | 3 ++- src/hermes-adapter.mjs | 11 +++++++---- src/model-routing.mjs | 3 ++- src/shadow-compaction.mjs | 14 +++----------- 4 files changed, 14 insertions(+), 17 deletions(-) diff --git a/src/browser.mjs b/src/browser.mjs index 668cf36..f63773b 100644 --- a/src/browser.mjs +++ b/src/browser.mjs @@ -125,6 +125,7 @@ export async function runBrowserFastPath({ harness = "browser", start_url = null, provider = "demo", + config, enabled, maxSteps = 8, maxSeconds = 30, @@ -169,7 +170,7 @@ export async function runBrowserFastPath({ if (elapsed(started) > maxSeconds * 1_000) return finish("handoff", "browser_fast_path_timeout", observation); let routed; try { - routed = await decideBrowserStep({ goal, observation, harness, start_url, progress }, { enabled: true, provider, policy }); + routed = await decideBrowserStep({ goal, observation, harness, start_url, progress }, { enabled: true, provider, config, policy }); } catch (error) { metrics.failures += 1; return finish("handoff", "browser_decision_failed", observation); diff --git a/src/hermes-adapter.mjs b/src/hermes-adapter.mjs index 90f929f..1616024 100644 --- a/src/hermes-adapter.mjs +++ b/src/hermes-adapter.mjs @@ -36,6 +36,7 @@ async function handle({ operation, args = {}, decision = null } = {}) { const { provider: _provider, engine = "native", ...request } = args; const result = await routeRequest(request, { provider, + config, engine, contextFilterMode: request.policy?.context_filter_mode ?? config.features?.context_filter, }); @@ -45,8 +46,9 @@ async function handle({ operation, args = {}, decision = null } = {}) { if (operation === "browser_step") { const { provider: _provider, enabled, ...input } = args; const routed = await decideBrowserStep(input, { - enabled: enabled ?? config.features?.browser_fast_path, - provider, + enabled: enabled ?? config.features?.browser_fast_path, + provider, + config, }); routed.decision.browser_action = routed.action ? { id: routed.action.id, operation: routed.action.operation, target_id: routed.action.target_id ?? null, @@ -59,6 +61,7 @@ async function handle({ operation, args = {}, decision = null } = {}) { const { provider: _provider, enabled, ...input } = args; return superviseWork({ ...input, + config, enabled: enabled ?? config.features?.supervision, provider, receiptPath: casesPath, @@ -66,7 +69,7 @@ async function handle({ operation, args = {}, decision = null } = {}) { } if (operation === "model_route") { const { provider: _provider, ...input } = args; - const result = await recommendModelRoute({ ...input, provider }); + const result = await recommendModelRoute({ ...input, provider, config }); await persistRoutingCase({ schema_version: 1, harness: input.harness ?? "hermes", @@ -86,7 +89,7 @@ async function handle({ operation, args = {}, decision = null } = {}) { } if (operation === "shadow_compaction") { const { provider: _provider, ...input } = args; - return buildShadowCompactionReport({ ...input, provider }); + return buildShadowCompactionReport({ ...input, provider, config }); } if (operation === "record_execution") { if (!decision || typeof decision !== "object") throw new TypeError("decision is required for record_execution"); diff --git a/src/model-routing.mjs b/src/model-routing.mjs index fdca461..fc63221 100644 --- a/src/model-routing.mjs +++ b/src/model-routing.mjs @@ -11,6 +11,7 @@ export async function recommendModelRoute({ harness = "unknown", policy = {}, provider = "demo", + config, } = {}) { if (!Array.isArray(models) || models.length === 0) throw new TypeError("models[] is required"); const profiles = models.map(normalizeProfile); @@ -25,7 +26,7 @@ export async function recommendModelRoute({ // A recommendation must never create an approval boundary by itself. confirmation_risk_levels: [], }, - }, { provider }); + }, { provider, config }); return { ...decision, route_mode: "shadow", diff --git a/src/shadow-compaction.mjs b/src/shadow-compaction.mjs index 6aaa8a1..193b365 100644 --- a/src/shadow-compaction.mjs +++ b/src/shadow-compaction.mjs @@ -1,7 +1,6 @@ import { sanitizeContext } from "./context-filter.mjs"; import { isPinnedEvidence } from "./relevance-filter.mjs"; -import { OpenRouterDecisionsProvider, TypeSafeProvider } from "./providers/typesafe.mjs"; -import { DemoProvider } from "./providers/demo.mjs"; +import { resolveProvider } from "./providers/index.mjs"; const LIST_KEYS = new Set(["messages", "events", "logs", "tool_results", "history", "transcript"]); const DEFAULT_BATCH_SIZE = 8; @@ -18,6 +17,7 @@ export async function buildShadowCompactionReport({ intent = "Assess which context records must remain available to a later agent.", context = {}, provider = "demo", + config, batchSize = DEFAULT_BATCH_SIZE, keepThreshold = DEFAULT_KEEP_THRESHOLD, minConfidence = DEFAULT_MIN_CONFIDENCE, @@ -38,7 +38,7 @@ export async function buildShadowCompactionReport({ let failed = null; try { - const resolved = resolveProvider(provider); + const resolved = resolveProvider(provider, { config }); if (!resolved || typeof resolved.evaluate !== "function") throw new Error("provider does not support batched evaluation"); for (const batch of batches(evaluable, normalizedBatchSize)) { const raw = await resolved.evaluate({ @@ -144,14 +144,6 @@ function readNoul(answer) { return { probability, confidence }; } -function resolveProvider(provider) { - if (provider && typeof provider === "object") return provider; - if (provider === "demo") return new DemoProvider(); - if (provider === "openrouter") return new OpenRouterDecisionsProvider(); - if (provider === "typesafe") return new TypeSafeProvider(); - throw new Error(`unsupported provider: ${provider}`); -} - function batches(items, size) { return Array.from({ length: Math.ceil(items.length / size) }, (_, index) => items.slice(index * size, (index + 1) * size)); } From 27dfa15e1939116c9a97d75b81f98119d86216cf Mon Sep 17 00:00:00 2001 From: typakon4 Date: Sat, 26 Sep 2026 00:55:52 +0600 Subject: [PATCH 3/3] docs: align README with main before merging feature work --- README.md | 127 ++++++++++++++++++++++-------------------------------- 1 file changed, 51 insertions(+), 76 deletions(-) diff --git a/README.md b/README.md index 46d2424..d1c2814 100644 --- a/README.md +++ b/README.md @@ -5,13 +5,11 @@

jev-layer

- Portable System-1 decision layer for agent harnesses.
- Host-owned routing, receipts, replay, and fail-open integrations. + Let an agent make a bounded choice without handing it control.
+ Jev recommends. Your harness still checks permissions and executes.

-

- Hermes · OMP · Codex · generic MCP -

+

Hermes · OMP · Codex · generic MCP

CI status @@ -22,68 +20,57 @@ [English](README.md) · [Русский](README.ru.md) · [简体中文](README.zh-CN.md) -jev-layer routes bounded choices and records evidence; the host keeps execution, permissions, approvals, retries, recovery, and final results. - -> Integrating jev-layer into a harness? Start with the [Agent implementation guide](docs/AGENT-IMPLEMENTATION.md), not this README alone. - -## Architecture +**Install:** `npm install --global jev-layer` · [Integrate a harness](docs/AGENT-IMPLEMENTATION.md) · [Security model](SECURITY.md) -

- Architecture: agent harnesses send bounded requests to jev-layer; the host owns permissions and execution; receipts support replay. -

+## What changes -Jev never executes a selected capability. A provider can be deterministic `demo`, OpenRouter Decisions, or TypeSafe; provider-backed tests are not required for normal CI. +| Without Jev | With Jev | +| --- | --- | +| Your harness follows its existing path to choose a capability. | The harness can ask `jev_route` to choose from a bounded set it supplies. | +| Your harness owns permissions, approvals, and execution. | Your harness still owns permissions, approvals, and execution. | +| Execution results stay in the host's normal workflow. | The host can attach the result to the decision with `jev_record_execution` and replay cases offline. | -## Quick Start +Jev never executes a selected capability. If it is disabled, unavailable, invalid, or inconclusive, control returns to the host's normal path. -Requirements: Node.js 20 or newer. There are no mandatory runtime dependencies. +## Quick start -Install the published CLI: +Requires Node.js 20 or newer. There are no mandatory runtime dependencies. ```sh npm install --global jev-layer -``` - -Or use a local clone: - -```sh -npm install -npm link jev install --project /path/to/workspace jev add generic --project /path/to/workspace jev doctor --project /path/to/workspace ``` -`npm link` is local only. It does not publish the package. Use `node /path/to/jev-layer/bin/jev.mjs ...` instead if a global link is not wanted. The default `demo` provider is offline and deterministic. - -To call the stdio MCP server directly: +The default `demo` provider is deterministic and works offline. To run the stdio MCP server directly: ```sh jev mcp ``` -To use a provider with credentials, keep keys outside the repository: +## How it fits into a harness -```sh -export JEV_LAYER_PROVIDER=openrouter -export OPENROUTER_API_KEY='provided-by-your-secret-store' -jev doctor --project /path/to/workspace -``` +

+ Architecture: agent harnesses send bounded requests to jev-layer; the host owns permissions and execution; receipts support replay. +

-All three modes (`demo`, `openrouter`, and direct `typesafe`), their endpoints, and configuration precedence are documented in the [provider guide](docs/PROVIDERS.md). +1. The host sends Jev a request and the candidate capabilities it already allows. +2. Jev returns a bounded recommendation. The host checks it against its own registry and permissions. +3. The host decides whether to execute, then can record what happened against the original `correlation_id`. -## Core surfaces +A provider can be deterministic `demo`, OpenRouter Decisions, or TypeSafe. Provider-backed tests are not required for normal CI; see the [provider guide](docs/PROVIDERS.md). -- **Routing:** `jev_route` selects one capability from the host-supplied candidate set. Selection is advisory; the host validates the id and permissions. -- **Receipts/replay:** `jev_record_execution` joins the host result to the original `correlation_id`. JSONL cases live in `.jev/replay/cases.jsonl` and can be evaluated offline with `npm run replay:evaluate`. -- **Supervision:** `jev_supervise` returns bounded work-state judgments; deterministic host policy maps them to `continue`, `verify`, `retry`, `finish`, or `escalate`. The receipt records a deterministic `evidence_state`: `present`, `missing`, or `contradictory`. Contradictory evidence cannot result in `finish`. Jev does not perform those actions. -- **Model routing:** `jev_model_route` returns one recommendation from host-declared model profiles for the next model call. It is shadow-only: the host must measure outcome, retries, latency, and cost before it changes any provider/model setting. Model-route decisions are joined to later `jev_record_execution` receipts by `correlation_id`; record `result.model_route = { actual_model_id, retry_count, outcome }`, then run `npm run model-route:report -- /path/to/cases.jsonl` for a read-only evidence report. -- **Shadow compaction:** `jev_shadow_compaction` makes batched, conservative, report-only keep/drop candidates for host-supplied context. It never mutates, summarizes, or deletes context; pinned paths, errors, commands, and requirements are retained without provider review, while provider failures retain every remaining item. +## What else it can do + +- **Supervision:** `jev_supervise` returns bounded work-state judgments. The host decides whether to continue, verify, retry, finish, or escalate. +- **Model routing:** `jev_model_route` recommends one host-declared model profile for a future call. It is advisory only; the host measures outcomes before changing provider or model settings. Correlated receipts can be reviewed with `npm run model-route:report -- /path/to/cases.jsonl`. +- **Shadow compaction:** `jev_shadow_compaction` produces report-only keep/drop candidates for host-supplied context. It does not summarize, mutate, or delete context, and keeps pinned evidence on provider failure. - **Context filtering:** optional deterministic `shadow` or `conservative` filtering reduces stale context without LLM summarization. -- **Experimental browser fast-path:** `jev_browser_step` chooses one bounded action from a host observation. The host supplies observations, approval, native execution, and recovery. It is opt-in and does not start a browser worker. -- **Fail-open:** disabled, unavailable, invalid, or inconclusive Jev calls return control to the host's normal path. Jev never widens permissions or guesses execution. +- **Experimental browser fast-path:** `jev_browser_step` recommends one bounded action from a host observation. The host supplies approval, native execution, and recovery; Jev does not start a browser worker. +- **Fail open:** optional Jev surfaces are disabled by default. Jev never widens permissions or guesses execution. -All optional surfaces are disabled by default: +Enable optional surfaces explicitly: ```sh JEV_BROWSER_FAST_PATH=1 jev mcp @@ -91,43 +78,35 @@ JEV_SUPERVISION=1 jev mcp JEV_CONTEXT_FILTER=shadow jev cli --input examples/route-request.json ``` -## Harness adapters - -Current examples live under `integrations/`: - -- `integrations/hermes/` — see the [local-agent handoff (Russian)](docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md) for the active plugin topology and safe change workflow. -- `integrations/omp/` -- `integrations/codex/` -- `integrations/template/` - -The release baseline records OMP `18.2.6`, Hermes `0.21.3` (`b675e6de`), and Codex CLI `0.155.1` observed in the preparation environment. This is a version/contract baseline, not a claim of full provider/model coverage; see [docs/COMPATIBILITY.md](docs/COMPATIBILITY.md). - -Adapters are intentionally thin. They may call the CLI or stdio MCP, but the host must retain native capability lookup, permissions, approvals, execution, retries, recovery, and final output. +For provider credentials, keep keys outside the repository: -### Adding a new harness +```sh +export JEV_LAYER_PROVIDER=openrouter +export OPENROUTER_API_KEY='provided-by-your-secret-store' +jev doctor --project /path/to/workspace +``` -The shortest PR path is: +## Pick an integration -1. copy `integrations/template/adapter.mjs`; -2. add `integrations//` and a secret-free config/example; -3. call `jev_route`, preserve `correlation_id`, execute only through the host registry, then call `jev_record_execution`; -4. add an offline smoke fixture for success, fail-open, approval denial, and execution receipt; -5. document supported versions and run CI. +Examples and adapters live under `integrations/`: -See [CONTRIBUTING.md](CONTRIBUTING.md) for the adapter contract and [docs/SCHEMA-VERSIONING.md](docs/SCHEMA-VERSIONING.md) for compatibility rules. +- Hermes: `integrations/hermes/` (see the [local-agent handoff guide](docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md)) +- OMP: `integrations/omp/` +- Codex: `integrations/codex/` +- Another harness: start from `integrations/template/` -## Browser status +Adapters stay thin. The host retains native capability lookup, permissions, approvals, execution, retries, recovery, and final output. For the adapter contract, see [CONTRIBUTING.md](CONTRIBUTING.md). For compatibility rules, see [docs/SCHEMA-VERSIONING.md](docs/SCHEMA-VERSIONING.md). -Browser fast-path **reliability is validated against the current real-browser fixtures**, including action sequencing, visible-link navigation, native select execution, approval denial, and recovery. **Performance optimization remains experimental**. No browser speedup claim is made. +The release baseline records OMP `18.2.6`, Hermes `0.21.3` (`b675e6de`), and Codex CLI `0.155.1` observed in the preparation environment. This is a version/contract baseline, not a claim of full provider/model coverage; see [docs/COMPATIBILITY.md](docs/COMPATIBILITY.md). -## Security and compatibility +## Boundaries -- MIT licensed; see [LICENSE](LICENSE). -- jev-layer is not a security boundary. Host permissions and approvals are authoritative; see [SECURITY.md](SECURITY.md). +- jev-layer is not a security boundary. Host permissions and approvals remain authoritative; see [SECURITY.md](SECURITY.md). - Schema, MCP tool, receipt, replay, and adapter contracts are currently version 1. Prefer additive changes; do not break v1 silently. +- Browser fast-path reliability is validated against current real-browser fixtures. Performance optimization remains experimental; no browser speedup claim is made. - Do not commit credentials, logs containing secrets, `.env` files, or machine-specific paths. -## Verification +## Verify locally ```sh npm test @@ -137,12 +116,8 @@ npm run clean-install-smoke npm pack --dry-run ``` -The GitHub Actions matrix runs these checks on Node.js 20, 22, and 24. Provider-backed tests require an explicitly configured secret-managed environment and are not part of ordinary PR CI. +GitHub Actions runs these checks on Node.js 20, 22, and 24. Provider-backed tests require a secret-managed environment and are not part of ordinary PR CI. -## Release documents +## Project docs -- [CONTRIBUTING.md](CONTRIBUTING.md) -- [SECURITY.md](SECURITY.md) -- [RELEASE.md](RELEASE.md) -- [CHANGELOG.md](CHANGELOG.md) -- [Agent implementation guide](docs/AGENT-IMPLEMENTATION.md) +[Agent implementation guide](docs/AGENT-IMPLEMENTATION.md) · [Providers](docs/PROVIDERS.md) · [Compatibility](docs/COMPATIBILITY.md) · [Contributing](CONTRIBUTING.md) · [Security](SECURITY.md) · [Release](RELEASE.md) · [Changelog](CHANGELOG.md)