diff --git a/README.md b/README.md index 8914ef8..d67814e 100644 --- a/README.md +++ b/README.md @@ -43,7 +43,8 @@ reported as success. There is no public-peer or OpenAI fallback. These first commands use the core directly. They are **not yet routed through Codex**. The separate app-server client implements the pinned NDJSON handshake, thread/turn requests, notifications and interruption, and declines tool approvals -until an interactive approval layer is connected. Its focused protocol tests are +by default. An explicit caller can supply a narrowly scoped per-command approval +policy; the normal extension does not enable it. Its focused protocol tests are now complemented by a **real, source-built app-server lifecycle trial**: initialization, an ephemeral VOLPAROSSA-provider thread, exact unsubscribe and clean shutdown pass in disposable namespaces without OpenAI credentials or @@ -58,6 +59,19 @@ a native-editor test or a measure of general coding quality. See the [explicit runtime build and native trial](docs/RUNTIME_BUILD.md). Nothing is downloaded or started merely by installing or activating the extension. +The next [local Responses adapter](docs/RESPONSES_PROVIDER.md) now connects a +bounded text/tool subset to the core's separate conversation interface. It retains +call/result identities and waits for confirmed core cleanup before returning a +completed turn. Its real HTTP/Unix-socket tests use synthetic model responses; +the actual Codex/model/tool loop is **not proved yet**. The new Qwen conversation +profile is a larger-context candidate, not evidence of reliable coding performance. + +An explicit [native coding trial](docs/NATIVE_CODING_TRIAL.md) now supplies the +missing model catalog and disposable read/edit/test harness. It uses the full +pinned Codex prompt, actual core inference and native tools, with approvals limited +to one synthetic project. The harness is implemented and its offline checks pass; +the actual model-driven coding trial is still pending. + ## Try the development extension On Linux, explicitly prepare and start the core's private service following its @@ -81,12 +95,12 @@ npm run check ## Remaining integration work -- Connect a genuine VOLPAROSSA coding-model/provider interface to the Codex - Responses/tool loop; the current bounded Q&A endpoint is not that interface. +- Prove the new conversation/provider interface with an actual model and the + native Codex Responses/tool loop; the bounded Q&A endpoint stays separate. - Connect the built runtime to the extension and core provider, retaining its isolated configuration and upstream notices; never use the owner's OpenAI login or cloud fallback. -- Add conversation/context support, typed tool calls, reviewable diffs and local +- Complete native conversation/tool interoperability, reviewable diffs and local approvals, then prove an actual edit-and-test coding task end to end. - Delegate eligible work through the core's cooperative scheduler, with explicit privacy scope, cancellation, resource accounting and result provenance. diff --git a/docs/NATIVE_CODING_TRIAL.md b/docs/NATIVE_CODING_TRIAL.md new file mode 100644 index 0000000..b74a62d --- /dev/null +++ b/docs/NATIVE_CODING_TRIAL.md @@ -0,0 +1,177 @@ +# Native Codex → core → real coding trial + +This additive development harness connects the **actual pinned app-server** to +the local Responses adapter and an already-running VOLPAROSSA Qwen conversation +service. The real trial now reaches the native runtime and model, but **a successful +native model-driven read/edit/test loop has not yet been observed**. Offline protocol/helper tests +are not a substitute for that result. + +The ordinary extension and `AppServer` defaults remain unchanged: no automatic +runtime/model startup, read-only threads and declined tool approvals. An explicit +caller may select an exact advertised model and supply a per-command approval +handler. Only a thread in that caller's exact writable root receives workspace +write access. The test harness uses that opt-in for one synthetic project; it +does not add an automatic-approval mode to normal editor use. + +## One small, real task + +The model receives the entire original 20,903-byte Codex base prompt from commit +`67727e7cf114cf3e1b71db368d74b24e32f6cb12`, SHA256 +`ac8ae107a0d72fe3476b430afb161ea4e67da2e446d778aefc44828160559807`. +The explicit `qwen3-0.6b-v1` model catalog chooses native `exec_command` / +`write_stdin` tools and no reasoning or grammar-based apply-patch tool. Unsupported +grammar remains an error in the provider, not something silently stripped from a +tool definition. There is no shortened replacement system prompt. + +The disposable project contains only `arithmetic.py`, whose `add` function is +incorrect. The model must choose actual native tool calls to: + +1. Read that file through the read-only mounted fixture helper. +2. Supply its own bounded arithmetic expression for the helper to write. +3. Run three actual unit tests and finish the native turn. + +The helper performs real I/O and tests. It does not generate a model response or +substitute a predetermined repair. Its three command forms are the only commands +the caller approves, bound to the exact thread, turn and working directory. The +caller declines session approvals, network requests, policy amendments, unknown +commands and file-change requests. Successful reads precede edits, and edits +precede tests. A separate post-turn test verifies the actual changed file again. + +This narrow task is an interoperability test, **not** evidence of general coding +ability, free-form shell authority, editor integration or private distributed +inference. The core remains responsible for model execution, limits and cleanup. +No peer scheduler or secondary inference engine is introduced here. + +## Explicit execution + +Requirements: Debian 13/Linux with unprivileged bubblewrap namespaces, an existing +verified source-built app-server bundle including `BUILD_REPORT.json` and notices, +a verified Node 22+ executable, the pinned upstream prompt, and an explicitly +started same-owner Qwen private-conversation service. The script fetches, builds +and starts none of these dependencies. Its core socket and parent must be mode +0600 and 0700 respectively. + +Choose a fresh output directory below this checkout's existing `build/` directory: + +```sh +python3 -B scripts/smoke_native_coding.py \ + --app-server /absolute/workspace/codex-runtime/runtime/codex-app-server \ + --app-server-sha256 VERIFIED_BINARY_SHA256 \ + --build-report /absolute/workspace/codex-runtime/BUILD_REPORT.json \ + --node /absolute/workspace/node --node-sha256 VERIFIED_NODE_SHA256 \ + --upstream-prompt /absolute/pinned-source/codex-rs/models-manager/prompt.md \ + --socket /absolute/owner-private-directory/private.sock \ + --output /absolute/checkout/build/new-native-coding-trial \ + --execute --yes +``` + +Without `--execute --yes`, only validation and the action plan run. Actual execution +uses new user, network, PID, mount, IPC and UTS namespaces with all capabilities +dropped. Only system runtime files, the exact socket, verified runtimes, adapter +sources and synthetic state are exposed. The owner's home and Codex configuration +are absent; neither `HOME` nor `CODEX_HOME` is overridden. The app-server uses a +new in-memory loopback bearer secret, not an OpenAI credential, and that secret is +excluded from tool environments. Outside networking, telemetry, external tools, +cloud login and automatic retries are disabled. + +The complete native task is bounded to 40 minutes and at most six accepted fixture +commands across at most two native turns. A normally completed turn is not proof +of a completed task: if observed read/edit/test actions are still missing, the +same thread receives at most one neutral continuation asking to finish the +original task. It does not supply an expression, command or answer. Failed, +interrupted or disconnected turns do not trigger continuation. Turn correlation, +approval ordering, the shared command budget and the original deadline remain in +force; neither budgets nor permissions reset for the second turn. +Every model request remains subject to the core's existing token, memory and +600-second execution bounds. The separate two-turn core KVM fixture currently +has a **1,400-second service window**, so this longer trial needs an explicitly +extended or separate service window; do not append it after that service stops. +Memory admission refusal, invalid model tools, token exhaustion or a missing +read/edit/test action fails the trial visibly. No reclaims, restarts, silent input +truncation or canned answers turn those failures into success. + +## Evidence and cleanup + +The original core [run 36913897403](https://github.com/VOLPAROSSA/volparossa/actions/runs/36913897403) +on `527e8ac35d9a0c0e461f76fac17e97ed10a93db9`, with Code +`7e35ba8d56df8ec43715119ceb0a1ae3f02f1f63`, successfully compiled and executed +the pinned native app-server and made two real private-model requests. One +response completed and one was incomplete; both confirmed worker cleanup. The +driver declined one command approval and accepted none, so no read/edit/test +action completed. The original closed receipt did not retain the rejection's +reason or private command text. It does not prove why approval failed or which +model-output condition caused the second incomplete result. + +Runtime exit was graceful, private state and services were cleaned up, no OOM +occurred and host network state was unchanged. The original eight-file artifact +SHA-256 is `0389a4754bd08470e815b2e452089ee933010711ca4bd3ddb02c1a6dbae002b4`. +Preserve this as a failed coding trial, not a successful loop or evidence of +general coding ability. + +A subsequent isolated, synthetic Responses protocol reproduction with the same +pinned native source observed a canonical read approval carrying +`proposedExecpolicyAmendment`. Every other exact fixture check matched; the +previous policy rejected the request solely because that proposal existed. +The local binary was `9635cc912ca720b1dd496ba34ca5be920ec46af4319d936a8aa5e59f57094e9f`, +not the CI binary. The protocol report SHA-256 is +`e58ea8b86b803821a1d7930e44a920c20ac2ccb2f12009c9da5ad252734af2dd`. +Every command in that reproduction was declined, the native runtime exited +cleanly, private state was removed and host network state was unchanged. This +is protocol evidence, not inference or command-execution evidence, and does +not reconstruct the original model's unretained approval payload. + +The correction treats the offered execpolicy rule as a proposal, not extra +authority. Exact command, working-directory, thread/turn, action-kind, network +and additional-permission checks remain. The adapter still emits only one-shot +`accept` or `decline`, never session or persistent-rule approval, consistent +with the [official app-server approval semantics](https://learn.chatgpt.com/docs/app-server#command-execution-approvals). +Receipt version 2 adds only fixed denial-category counts (including ordering +and command budget), whose sum equals declined approvals; it never retains +commands, arguments, paths, identifiers or stderr. Historical version-1 receipts +remain readable. A fresh actual model-driven coding trial remains required. +The same isolated native protocol probe now evaluates the observed request as +authorized under the corrected policy, while still returning `decline` for the +probe itself: zero commands executed, clean native exit and unchanged host state. +That after-fix protocol report SHA-256 is +`1ba245d00f749c66ba1dd4af5500420d9be05c078d21f2db8fb8c57ba0d7abe6`. + +The original [run 36925945880](https://github.com/VOLPAROSSA/volparossa/actions/runs/36925945880) +on core `ff2abe632a6779494301685014f89ec49aa2d259` / Code +`eb48696eb37afb9cda59bffc350845309b963dbb` now completes an actual native read: +one approval accepted, none declined, two real model responses completed and both +worker cleanups confirmed. The native turn ends normally **without editing or +testing**, so the task fails its actual-action assertion, before independent tests. +Elapsed time is 790,763 ms and service CPU usage 1,559,652,831 microseconds; +peak memory is 1,883,226,112 bytes with no OOM. No timeout or forced stop occurs, +private/runtime cleanup passes and host state is unchanged. The original eight-file +artifact SHA-256 is `6aafeaee20a247d05f0e334bad07ac630e810417719783f8c8d45138ac80abe4`. +The second response's text was not retained; these counters do not establish why +the model stopped or whether it attempted a tool in an unsupported text format. + +Receipt version 3 therefore adds opt-in, closed per-response diagnostics: output +kind, prompt/generated token counts, completion/incomplete reason and elapsed +time from submission through confirmed cleanup, at most 16 records. Fixed native +completed-item type counters and started/completed turn counts distinguish model +text, tool proposals and actual command events. They never retain text, commands, +arguments, paths or identifiers. Historical version-1/2 receipts remain readable. +The bounded continuation and these diagnostics have offline controller/protocol +coverage, not a newly successful model-driven read/edit/test proof. + +Success requires the actual app-server's command-completion events, changed file +hash, independent passing tests, at least four cleanup-confirmed real core +responses, exact thread unsubscribe and graceful runtime exit. Partial responses, +unexpected commands and uncertain cleanup cannot pass. The JSON report contains +closed statuses, counters and source/runtime/input hashes—not prompts, model text, +tool output, URLs or credentials. Runtime source/lock/patch provenance is checked +against the tracked build pin, and the original upstream notices remain intact. + +The owned PID namespace is joined and the synthetic project, ephemeral app-server +home and temporary configuration are removed. Read-only before/after snapshots +must show unchanged host network namespace, routes and DNS. A surrounding KVM +fixture must independently account for its actual core/model service lifecycle. + +Offline checks require no model or app-server execution: + +```sh +node --test tests/app-server.test.cjs tests/native-coding.test.cjs tests/responses-provider.test.cjs +``` diff --git a/docs/RESPONSES_PROVIDER.md b/docs/RESPONSES_PROVIDER.md new file mode 100644 index 0000000..6204caa --- /dev/null +++ b/docs/RESPONSES_PROVIDER.md @@ -0,0 +1,134 @@ +# Local Responses → core conversation adapter + +This executable adapter connects the open Codex runtime's HTTP Responses wire to +the **existing VOLPAROSSA private-conversation service**. It implements no model, +peer scheduler or tool executor. It is not a decentralized/private-peer inference +claim: this core operation explicitly advertises `private_local`, no network, +public cache, training or cloud fallback. + +`src/private-conversation.cjs` negotiates the additive conversation handshake on +the same-owner, mode-0600 Unix socket. It reuses the original Q&A connection, +framing reader, owner checks and cancellation lifecycle; `private-compute.cjs` +is unchanged. Conversation-only request framing may use the separately negotiated +larger bound. No legacy Q&A request or response is rewritten. + +## Explicit application-owned startup + +Requiring the module or opening an editor starts nothing. An integrating local +application explicitly calls: + +```javascript +const { startResponsesProvider } = require('../src/responses-provider.cjs'); +const provider = await startResponsesProvider({ + socketPath: '/absolute/owner-private-directory/compute.sock', + model: 'qwen3-0.6b-v1', +}); +// Pass provider.baseUrl and provider.bearerToken directly to the explicitly +// launched local runtime. Keep the random token in memory, never in logs or Git. +// On application shutdown: +await provider.close(); +``` + +The service binds only an ephemeral `127.0.0.1` port, exposes `POST /v1/responses`, +and requires its newly generated 256-bit bearer secret. This is **not an OpenAI +credential**. Host-header checks and rejection of Origin/Referer block browser +cross-origin use; there is no CORS, WebSocket, remote-listener or cloud route. +Same-user processes are not isolated adversaries. The adapter does not modify +Codex authentication, global configuration, user home or editor settings. + +The endpoint permits one active request, eight connections, bounded headers and +body-read time, and at most 512 KiB request bodies. Actual model-specific limits +are narrower and checked without truncation. HTTP disconnect/shutdown requests +the exact core task's cancellation. The slot remains occupied until terminal +core cleanup or a bounded cleanup-uncertainty failure; an acknowledgement alone +is never presented as successful cleanup. No automatic retry occurs. + +## Supported, lossless subset + +Requests select the exact operator model, `store:false`, `stream:true`, and +`tool_choice:"auto"` when supplied. Instructions, ordered text-only messages, +function/custom calls and their correlated tool results are translated to the +core's typed history. Plain text parts are joined in their original order with +two newline separators; images, audio, binary attachments and unsupported item +types are refused. Original call IDs and explicit namespaces survive round trips. +Every historical call must refer to an offered tool and have exactly one matching +result before a later message. The adapter never silently drops old tools. + +Function tools pass their object schema to the core as data; neither layer claims +arbitrary JSON Schema enforcement. `strict:true` is rejected. Literal custom +tools are supported; grammar formats are rejected rather than weakened. Namespace +descriptions are retained in each member's bounded description. A complete core +turn proposes at most one tool call, even when the request permits parallel calls. +The native tool harness still owns argument validation, approvals and execution. +The separate [native coding trial](NATIVE_CODING_TRIAL.md) supplies an exact Qwen +catalog and an explicit synthetic-workspace approval policy. Its implementation +does not by itself prove a successful native tool loop. + +The exact pinned Codex builder includes several optional transport fields even +for external providers. The adapter accepts these explicitly: + +- `reasoning:{}` or no-thinking settings (`effort:"none"`, `summary:"none"`, + `context:"current_turn"`); actual reasoning requests are refused. +- `include:["reasoning.encrypted_content"]`: there are no reasoning items to + expand, so none are fabricated. +- Bounded `prompt_cache_key` and string-valued `client_metadata`: discarded in + memory, never sent to the model, stored, logged or used as authority. No cache + hit or telemetry service is claimed. +- Empty text controls, default service tier and optional usage-stream request. + +Stateful response IDs, built-in hosted tools, structured-output constraints, +unknown options, ambiguous duplicate JSON keys and unsafe integer values are +explicit errors. No prompt, schema, tool result or model output is truncated. + +## Model budgets and truthful completion + +The SmolLM2 profiles remain unchanged: 135M uses 192 prompt / 64 output tokens; +360M and 1.7B use 1,024 / 256. **A normal Codex prompt does not fit these profiles.** +The core tokenizer remains authoritative and rejects oversize input. + +The new explicit `qwen3-0.6b-v1` conversation contract reserves 12,288 prompt and +1,024 output tokens within a conservative 32,768-token model context. Its typed +input bound is 256 KiB, instructions 64 KiB, history 128 items and tools 32; +messages are at most 64 KiB, descriptions 8 KiB, and individual tool payloads +remain 4 KiB. Its conversation-only request frame is 512 KiB; response frames +remain 64 KiB and model output 4 KiB. These are advertised limits, not measured +memory usage or evidence that arbitrary projects fit. The adapter accepts the +exact core limits and does not raise them. Qwen history preserves system/developer +roles and ordering; Smol still rejects those roles. + +Only a validated terminal core result with confirmed worker/staging cleanup can +produce SSE. Complete assistant/tool data emits the corresponding ordered item, +text/argument events and `response.completed`. This is buffered protocol +adaptation, **not** backend token-by-token streaming. `token_limit`, truncated +wire data and invalid generated syntax instead produce `response.incomplete` +with no executable output and never `response.completed`. A syntactically complete +tool proposal is not proof that the proposal is safe, correct or accomplished. + +## Verification and source binding + +Focused tests use real loopback HTTP and same-owner Unix framing with a clearly +synthetic protocol peer. They cover exact text/tool round trips, native-shaped +request metadata, a Qwen frame larger than 32 KiB without truncation, owner and +capability gates, cleanup-before-SSE, cancellation and incomplete/error handling. +The unchanged Q&A regression tests pass alongside them. These checks do **not** +execute a model, native Codex edit/test loop, editor integration or private peers. + +```sh +node --test tests/private-conversation.test.cjs tests/responses-provider.test.cjs tests/private-compute.test.cjs +``` + +Interoperability targets the unmodified open Codex source +`67727e7cf114cf3e1b71db368d74b24e32f6cb12`: request construction in +`codex-rs/core/src/client.rs`, request schemas in `codex-rs/codex-api/src/common.rs`, +tool types in `codex-rs/tools/src/responses_api.rs`, and SSE consumption in +`codex-rs/codex-api/src/sse/responses.rs`. That parser treats incomplete reasons +other than `interrupted` as stream failures; this adapter never mislabels model +exhaustion as a successful interrupted response. The core's +`crates/volparossa/src/compute/private_conversation/WIRE.md` remains the authoritative +IPC contract. This is independently written GPL-3.0-only protocol glue; no upstream +source is copied here. Existing runtime provenance and notices remain unchanged. + +The event/tool shapes were also checked against the official +[Responses streaming documentation](https://developers.openai.com/api/docs/guides/streaming-responses) +and [function/custom-tool documentation](https://developers.openai.com/api/docs/guides/function-calling). +Only the explicitly described subset is implemented, not the entire hosted API. diff --git a/scripts/build_codex_runtime.py b/scripts/build_codex_runtime.py index 818205d..e384bb3 100644 --- a/scripts/build_codex_runtime.py +++ b/scripts/build_codex_runtime.py @@ -191,21 +191,33 @@ def size_bound(state): return total -def run_step(name, argv, env, state, *, network=False, seconds=3600, json_output=False): +def sandbox_command(argv, state, *, network=False): source = state / 'source' command = ['/usr/bin/bwrap', '--die-with-parent', '--ro-bind', '/', '/', '--tmpfs', '/home', '--tmpfs', '/root', '--tmpfs', '/run', '--tmpfs', '/tmp', '--bind', str(state), str(state), '--ro-bind', str(source), str(source), '--proc', '/proc', '--dev', '/dev', '--chdir', str(source / 'codex-rs')] - if not network: + if network: + # Resolve before /run is hidden: Ubuntu points /etc/resolv.conf into + # /run/systemd/resolve. Re-expose only this file, never the runtime tree. + resolver = Path('/etc/resolv.conf').resolve(strict=True) + require(resolver.is_file() and resolver.stat().st_size <= 64 * 1024, + 'invalid dependency-fetch resolver file') + command += ['--ro-bind', str(resolver), str(resolver)] + else: command += ['--unshare-net'] - command += ['--'] + argv + return command + ['--'] + argv + + +def run_step(name, argv, env, state, *, network=False, seconds=3600, json_output=False): + command = sandbox_command(argv, state, network=network) stamp = time.time_ns() log = state / f'{name}-{stamp}.log' stdout = state / f'{name}-{stamp}.stdout.json' if json_output else log receipt = state / f'{name}-{stamp}.json' report = dict(version=1, step=name, argv=argv, passed=False, network_enabled=network, source_read_only=True, compiler_jobs=2, private_home_hidden=True, + resolver_file_read_only=network, app_server_executed=False, log=log.name) process, start = None, time.monotonic() try: diff --git a/scripts/native_coding_fixture.py b/scripts/native_coding_fixture.py new file mode 100644 index 0000000..b068150 --- /dev/null +++ b/scripts/native_coding_fixture.py @@ -0,0 +1,87 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: GPL-3.0-only +"""Three real bounded actions on a disposable synthetic file; no model or canned repair.""" +import ast +import hashlib +import json +import os +from pathlib import Path +import stat +import sys +import unittest + +PROJECT = Path('/opt/work/project') +SOURCE = PROJECT / 'arithmetic.py' +ORIGINAL = 'def add(a, b):\n return a - b\n' + + +def source_text(): + info = SOURCE.lstat() + if not stat.S_ISREG(info.st_mode) or info.st_uid != os.getuid() or info.st_nlink != 1 or info.st_size > 256: + raise ValueError('fixture_source') + return SOURCE.read_text(encoding='ascii') + + +def expression(value): + if not 1 <= len(value) <= 80: + raise ValueError('expression_bound') + tree = ast.parse(value, mode='eval') + nodes = list(ast.walk(tree)) + if len(nodes) > 24 or any(type(node) not in (ast.Expression, ast.BinOp, ast.UnaryOp, ast.Name, + ast.Load, ast.Constant, ast.Add, ast.Sub, ast.Mult, ast.Div, ast.UAdd, ast.USub) for node in nodes): + raise ValueError('expression_scope') + for node in nodes: + if isinstance(node, ast.Name) and node.id not in ('a', 'b'): + raise ValueError('expression_name') + if isinstance(node, ast.Constant) and (type(node.value) is not int or not -100 <= node.value <= 100): + raise ValueError('expression_number') + return value + + +def checked_source(value): + prefix = 'def add(a, b):\n return ' + if not value.startswith(prefix) or not value.endswith('\n') or '\n' in value[len(prefix):-1]: + raise ValueError('source_shape') + expression(value[len(prefix):-1]) + return value + + +def main(): + if Path.cwd() != PROJECT or os.getuid() == 0: + raise ValueError('fixture_scope') + current = checked_source(source_text()) + action = sys.argv[1] if len(sys.argv) >= 2 else '' + if action == 'read' and len(sys.argv) == 2: + print(json.dumps(dict(action='read', source=current))) + elif action == 'edit' and len(sys.argv) == 3: + replacement = 'def add(a, b):\n return ' + expression(sys.argv[2]) + '\n' + # Validated exact fixture inode; no arbitrary filename or executable input. + fd = os.open(SOURCE, os.O_WRONLY | os.O_TRUNC | os.O_NOFOLLOW) + with os.fdopen(fd, 'w', encoding='ascii') as output: + output.write(replacement) + print(json.dumps(dict(action='edit', sha256=hashlib.sha256(replacement.encode()).hexdigest()))) + elif action == 'test' and len(sys.argv) == 2: + namespace = {} + exec(compile(current, str(SOURCE), 'exec'), {'__builtins__': {}}, namespace) + add = namespace['add'] + class ArithmeticTests(unittest.TestCase): + def test_positive(self): + self.assertEqual(add(2, 3), 5) + def test_zero(self): + self.assertEqual(add(9, 0), 9) + def test_negative(self): + self.assertEqual(add(-4, 2), -2) + result = unittest.TextTestRunner(verbosity=0).run(unittest.defaultTestLoader.loadTestsFromTestCase(ArithmeticTests)) + print(json.dumps(dict(action='test', passed=result.wasSuccessful(), tests=result.testsRun))) + return 0 if result.wasSuccessful() else 1 + else: + raise ValueError('fixture_action') + return 0 + + +if __name__ == '__main__': + try: + raise SystemExit(main()) + except (OSError, ValueError, SyntaxError, IndexError, ZeroDivisionError): + print('synthetic_fixture_action_failed', file=sys.stderr) + raise SystemExit(1) from None diff --git a/scripts/smoke_native_coding.cjs b/scripts/smoke_native_coding.cjs new file mode 100644 index 0000000..c0f4601 --- /dev/null +++ b/scripts/smoke_native_coding.cjs @@ -0,0 +1,125 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Actual pinned app-server + actual core conversation. No synthetic model reply. +'use strict'; +const fs = require('node:fs'); +const {createHash} = require('node:crypto'); +const {spawn, spawnSync} = require('node:child_process'); +const assert = require('node:assert/strict'); +const {AppServer} = require('/opt/src/app-server.cjs'); +const {PrivateConversation} = require('/opt/src/private-conversation.cjs'); +const {startResponsesProvider} = require('/opt/src/responses-provider.cjs'); +const {MODEL, PROJECT, modelCatalog, runtimeSettings, APPROVAL_DENIALS, + ITEM_TYPES, NativeTaskController} = require('/opt/src/native-coding-fixture.cjs'); +const sha = data => createHash('sha256').update(data).digest('hex'); + +async function main() { + assert.equal(process.cwd(), PROJECT); + assert.notEqual(process.getuid(), 0); + for (const key of ['HOME', 'CODEX_HOME', 'OPENAI_API_KEY', 'OPENAI_BASE_URL']) assert.equal(process.env[key], undefined); + const original = fs.readFileSync(`${PROJECT}/arithmetic.py`); + const instructions = fs.readFileSync('/opt/upstream-prompt.md', 'utf8'); + const report = {version: 3, kind: 'native-codex-core-coding', success: false, phase: 'capabilities', + model: MODEL, full_native_prompt_sha256: sha(instructions), before_sha256: sha(original), after_sha256: null, + native_turn_completed: false, read: false, edit: false, test: false, independent_test_passed: false, + accepted_commands: 0, declined_commands: 0, + approval_denials: Object.fromEntries(APPROVAL_DENIALS.map(reason => [reason, 0])), + turns_started: 0, turns_completed: 0, + item_types: Object.fromEntries(ITEM_TYPES.map(type => [type, 0])), response_diagnostics: null, + unexpected_command: false, thread_unsubscribed: false, + responses: null, private_peer_execution_claimed: false, general_coding_quality_claimed: false, + runtime_exit: null, forced_stop: false, diagnostic: null}; + let provider, child, client, exited, threadId, controller; + let stderrBytes = 0; + let deadline; + async function stop() { + if (controller) await controller.stop(); + else client?.close(); + } + try { + const preflight = new PrivateConversation('/opt/core/compute.sock'); + try { + const caps = await preflight.connect(); + assert.equal(caps.model_profile, MODEL); + assert.equal(caps.native_tool_template, true); + assert.equal(caps.local_only, true); + } finally { preflight.close(); } + fs.writeFileSync('/opt/catalog.json', JSON.stringify(modelCatalog(instructions)), {flag: 'wx', mode: 0o444}); + provider = await startResponsesProvider({socketPath: '/opt/core/compute.sock', model: MODEL, diagnostics: true}); + report.phase = 'launch'; + child = spawn('/opt/codex-app-server', ['--listen', 'stdio://', '--strict-config', + ...runtimeSettings(provider.baseUrl).flatMap(value => ['-c', value])], { + stdio: ['pipe', 'pipe', 'pipe'], env: {...process.env, VOLPAROSSA_PROVIDER_TOKEN: provider.bearerToken}}); + exited = new Promise(resolve => { + child.once('error', () => { controller?.closed(); resolve({code: null, signal: 'spawn_failed'}); }); + child.once('close', (code, signal) => { controller?.closed(); resolve({code, signal}); }); + }); + child.stderr.on('data', data => { + stderrBytes += data.length; + if (stderrBytes > 1024 * 1024) { report.diagnostic = 'stderr_bound'; void stop(); } + }); // Never store stderr, bearer tokens, model prompts, tool arguments or raw output. + client = new AppServer(child.stdout, child.stdin, {timeoutMs: 30000, allowedModels: [MODEL], writableRoot: PROJECT, + commandApproval(params) { + if (controller) return controller.approve(params); + report.declined_commands++; report.approval_denials.lineage++; + if (report.declined_commands > 2) void stop(); + return false; + }}); + report.phase = 'initialize'; await client.initialize(); + report.phase = 'thread-start'; + const started = await client.startThread({model: MODEL, cwd: PROJECT}); + assert.equal(started.model, MODEL); assert.equal(started.thread.modelProvider, 'volparossa'); + assert.equal(started.thread.ephemeral, true); assert.equal(started.cwd, PROJECT); + threadId = started.thread.id; + controller = new NativeTaskController(client, threadId, report); + report.phase = 'native-turn'; + deadline = setTimeout(() => { report.diagnostic = 'turn_deadline'; void stop(); }, 2400000); + await controller.run(); + assert(report.read && report.edit && report.test && !report.unexpected_command); + report.phase = 'independent-check'; + const checked = spawnSync('/usr/bin/python3', ['-B', '/opt/fixture.py', 'test'], { + cwd: PROJECT, encoding: 'utf8', maxBuffer: 8192, timeout: 5000}); + assert.equal(checked.status, 0); + assert.deepEqual(JSON.parse(checked.stdout), {action: 'test', passed: true, tests: 3}); + report.independent_test_passed = true; + report.after_sha256 = sha(fs.readFileSync(`${PROJECT}/arithmetic.py`)); + assert.notEqual(report.after_sha256, report.before_sha256); + assert.deepEqual(fs.readdirSync(PROJECT).sort(), ['arithmetic.py']); + report.responses = provider.observations; + assert(report.responses.completed >= 4 && report.responses.incomplete === 0 && + report.responses.submitted === report.responses.cleanup_confirmed); + report.phase = 'unsubscribe'; + assert.equal((await client.request('thread/unsubscribe', {threadId})).status, 'unsubscribed'); + report.thread_unsubscribed = true; + } catch { + report.diagnostic ??= 'native_coding_incomplete'; + } finally { + clearTimeout(deadline); + if (!report.native_turn_completed) await stop(); + controller?.dispose(); + client?.close(); + if (provider) { + try { await provider.close(); } + catch { report.diagnostic = 'provider_cleanup_unconfirmed'; } + report.responses = provider.observations; + report.response_diagnostics = provider.diagnostics; + } + if (child) { + let timer; + let exit = await Promise.race([exited, new Promise(resolve => { timer = setTimeout(() => resolve(null), 10000); })]); + clearTimeout(timer); + if (!exit) { + report.forced_stop = true; child.kill('SIGTERM'); + const hard = setTimeout(() => child.kill('SIGKILL'), 5000); + exit = await exited; clearTimeout(hard); + } + report.runtime_exit = exit.code; + } + report.success = report.native_turn_completed && report.read && report.edit && report.test && + report.independent_test_passed && report.thread_unsubscribed && !report.unexpected_command && + !report.diagnostic && !report.forced_stop && report.runtime_exit === 0; + if (report.success) report.phase = 'complete'; + fs.writeFileSync('/opt/work/receipt.json', JSON.stringify(report) + '\n', {flag: 'wx', mode: 0o600}); + if (!report.success) process.exitCode = 1; + } +} +main().catch(() => { process.exitCode = 1; }); diff --git a/scripts/smoke_native_coding.py b/scripts/smoke_native_coding.py new file mode 100644 index 0000000..0b204a7 --- /dev/null +++ b/scripts/smoke_native_coding.py @@ -0,0 +1,332 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: GPL-3.0-only +"""Explicit native Codex/core coding trial; no downloads or service startup.""" +import argparse +import hashlib +import json +import os +from pathlib import Path +import pwd +import re +import shutil +import signal +import socket +import stat +import subprocess +import sys +import tempfile + +ROOT = Path(__file__).resolve().parents[1] +UPSTREAM = '67727e7cf114cf3e1b71db368d74b24e32f6cb12' +PROMPT_SHA256 = 'ac8ae107a0d72fe3476b430afb161ea4e67da2e446d778aefc44828160559807' +ORIGINAL = b'def add(a, b):\n return a - b\n' +SOURCES = ('app-server.cjs', 'private-compute.cjs', 'private-conversation.cjs', + 'responses-provider.cjs', 'native-coding-fixture.cjs') +PHASES = ('capabilities', 'launch', 'initialize', 'thread-start', 'native-turn', + 'independent-check', 'unsubscribe', 'complete') +DIAGNOSTICS = (None, 'stderr_bound', 'turn_deadline', 'native_coding_incomplete', + 'provider_cleanup_unconfirmed') +APPROVAL_DENIALS = ('lineage', 'kind', 'item', 'cwd', 'command', 'network', + 'permissions', 'network_policy', 'order', 'budget') + + +def require(value, code): + if not value: + raise ValueError(code) + + +def digest(path): + with path.open('rb') as source: + return hashlib.file_digest(source, 'sha256').hexdigest() + + +def verified_file(value, expected, executable=False): + path = Path(value) + require(path.is_absolute() and path.resolve(strict=True) == path, 'canonical-file-required') + info = path.lstat() + require(stat.S_ISREG(info.st_mode) + and (not executable or not info.st_mode & 0o6022 and os.access(path, os.X_OK)) + and re.fullmatch('[0-9a-f]{64}', expected) and digest(path) == expected, + 'file-hash-or-mode') + return path + + +def private_socket(value): + path = Path(value) + require(path.is_absolute() and path.resolve(strict=True) == path, 'canonical-socket-required') + info, parent = path.lstat(), path.parent.lstat() + require(stat.S_ISSOCK(info.st_mode) and info.st_uid == os.getuid() + and stat.S_IMODE(info.st_mode) == 0o600 and stat.S_ISDIR(parent.st_mode) + and parent.st_uid == os.getuid() and stat.S_IMODE(parent.st_mode) == 0o700, + 'private-socket-required') + return path + + +def verified_build(value, binary, binary_hash): + path = Path(value) + require(path.is_absolute() and path.resolve(strict=True) == path and path.is_file() + and path.stat().st_size <= 65536, 'build-report-required') + report = json.loads(path.read_text()) + pin = json.loads((ROOT / 'third_party/codex-runtime.json').read_text()) + require(report.get('version') == 1 and report.get('app_server_built') is True + and report.get('staged_source_verified') is True and report.get('original_source_unchanged') is True + and report.get('source_revision') == pin['revision'] == UPSTREAM + and report.get('source_tree') == pin['tree'] + and report.get('lock_sha256') == pin['sha256']['codex-rs/Cargo.lock'] + and report.get('local_patches') == pin['patches'] + and report.get('binary') == dict(path='runtime/codex-app-server', + bytes=binary.stat().st_size, sha256=binary_hash) + and binary == path.parent / 'runtime/codex-app-server', 'build-source-binding') + notices = {name: digest(verified_file(str(path.parent / 'notices' / name), pin['sha256'][name])) + for name in ('LICENSE', 'NOTICE')} + return dict(report_sha256=digest(path), source_revision=pin['revision'], source_tree=pin['tree'], + lock_sha256=report['lock_sha256'], local_patches=pin['patches'], notices=notices) + + +def snapshot(): + return dict(network_namespace=os.readlink('/proc/self/ns/net'), + routes=digest(Path('/proc/net/route')), routes6=digest(Path('/proc/net/ipv6_route')), + dns=digest(Path('/etc/resolv.conf'))) + + +def command(binary, node, ipc, prompt, work, home): + # Only system runtime files, exact owned inputs and synthetic state are exposed. + # No broad root/home/workspace bind and no HOME/CODEX_HOME override. + result = ['/usr/bin/bwrap', '--die-with-parent', '--new-session', '--unshare-user', + '--uid', str(os.getuid()), '--gid', str(os.getgid()), '--unshare-net', '--unshare-pid', + '--unshare-ipc', '--unshare-uts', '--cap-drop', 'ALL', '--ro-bind', '/usr', '/usr', + '--symlink', 'usr/bin', '/bin', '--symlink', 'usr/sbin', '/sbin', + '--symlink', 'usr/lib', '/lib', '--symlink', 'usr/lib64', '/lib64', + '--dir', '/etc', '--ro-bind', '/etc/passwd', '/etc/passwd', + '--ro-bind', '/etc/group', '/etc/group', '--ro-bind', '/etc/ld.so.cache', '/etc/ld.so.cache', + '--tmpfs', '/tmp', '--dir', '/run', '--dir', '/opt', '--dir', '/opt/src', + '--dir', home, '--proc', '/proc', '--dev', '/dev', + '--ro-bind', str(binary), '/opt/codex-app-server', '--ro-bind', str(node), '/opt/node', + '--ro-bind', str(prompt), '/opt/upstream-prompt.md', + '--ro-bind', str(work / 'ipc'), '/opt/core', '--ro-bind', str(ipc), '/opt/core/compute.sock', + '--bind', str(work / 'state'), '/opt/work'] + for source in SOURCES: + result += ['--ro-bind', str(ROOT / 'src' / source), '/opt/src/' + source] + for source, destination in (('smoke_native_coding.cjs', '/opt/smoke_native_coding.cjs'), + ('native_coding_fixture.py', '/opt/fixture.py'), ('smoke_native_coding.py', '/opt/runner.py')): + result += ['--ro-bind', str(ROOT / 'scripts' / source), destination] + return result + ['--chdir', '/opt/work/project', '--clearenv', '--setenv', 'PATH', '/usr/bin:/bin', + '--setenv', 'LANG', 'C.UTF-8', '--', '/usr/bin/python3', '-B', '/opt/runner.py', + '--inside', os.readlink('/proc/self/ns/net')] + + +def inside(parent): + require(os.geteuid() != 0 and os.readlink('/proc/self/ns/net') != parent, 'isolation') + require({name for _, name in socket.if_nameindex()} <= {'lo'}, 'network-isolation') + caps = next(line.split()[1] for line in Path('/proc/self/status').read_text().splitlines() + if line.startswith('CapEff:')) + require(int(caps, 16) == 0, 'capability-isolation') + require(not any(name in os.environ for name in ('HOME', 'CODEX_HOME', 'OPENAI_API_KEY', 'OPENAI_BASE_URL')), + 'clean-environment') + account_home = Path(pwd.getpwuid(os.getuid()).pw_dir) + require(account_home.is_dir() and not list(account_home.iterdir()), 'empty-isolated-home') + private_socket('/opt/core/compute.sock') + return subprocess.run(['/opt/node', '/opt/smoke_native_coding.cjs'], check=False, timeout=2470).returncode + + +def check_native_diagnostics(value): + for key in ('turns_started', 'turns_completed'): + require(type(value[key]) is int and 0 <= value[key] <= 2, 'receipt-turn-count') + require(value['turns_completed'] <= value['turns_started'], 'receipt-turn-order') + counts = value['item_types'] + require(type(counts) is dict and set(counts) == {'commandExecution', 'agentMessage', 'userMessage', 'reasoning', 'other'} + and all(type(n) is int and 0 <= n <= 64 for n in counts.values()) + and sum(counts.values()) <= 64, 'receipt-item-counts') + diagnostic = value['response_diagnostics'] + if diagnostic is None: + require(value['responses'] is None, 'receipt-diagnostics-missing') + return + require(type(diagnostic) is dict and set(diagnostic) == {'version', 'records', 'truncated'} + and type(diagnostic['version']) is int and diagnostic['version'] == 1 + and type(diagnostic['truncated']) is bool and type(diagnostic['records']) is list + and len(diagnostic['records']) <= 16, 'receipt-response-diagnostics') + for row in diagnostic['records']: + require(type(row) is dict and set(row) == {'output_kind', 'prompt_tokens', 'generated_tokens', + 'turn_complete', 'incomplete_reason', 'elapsed_ms'} + and row['output_kind'] in ('assistant', 'function_call', 'custom_tool_call', 'incomplete') + and type(row['turn_complete']) is bool + and row['turn_complete'] == (row['output_kind'] != 'incomplete') + and (row['incomplete_reason'] is None if row['turn_complete'] else + row['incomplete_reason'] in ('token_limit', 'wire_truncated', 'invalid_output')), + 'receipt-response-shape') + for key, minimum, maximum in (('prompt_tokens', 1, 12288), ('generated_tokens', 1, 1024), + ('elapsed_ms', 0, 3600000)): + require(type(row[key]) is int and minimum <= row[key] <= maximum, 'receipt-response-bound') + counters = value['responses'] + require(counters is not None and len(diagnostic['records']) == min(16, counters['cleanup_confirmed']) + and diagnostic['truncated'] == (counters['cleanup_confirmed'] > 16), 'receipt-response-coverage') + if not diagnostic['truncated']: + require(sum(row['turn_complete'] for row in diagnostic['records']) == counters['completed'] + and sum(not row['turn_complete'] for row in diagnostic['records']) == counters['incomplete'], + 'receipt-response-correlation') + + +def closed_receipt(path): + require(path.is_file() and not path.is_symlink() and path.stat().st_size <= 16384, 'receipt-bound') + value = json.loads(path.read_text()) + booleans = ('success', 'native_turn_completed', 'read', 'edit', 'test', 'independent_test_passed', + 'unexpected_command', 'thread_unsubscribed', 'private_peer_execution_claimed', + 'general_coding_quality_claimed', 'forced_stop') + keys = set(booleans) | {'version', 'kind', 'phase', 'model', 'full_native_prompt_sha256', + 'before_sha256', 'after_sha256', 'accepted_commands', 'declined_commands', 'responses', + 'runtime_exit', 'diagnostic'} + require(isinstance(value, dict) and type(value.get('version')) is int + and value['version'] in (1, 2, 3), 'receipt-version') + if value['version'] >= 2: + keys.add('approval_denials') + if value['version'] == 3: + keys.update(('turns_started', 'turns_completed', 'item_types', 'response_diagnostics')) + require(isinstance(value, dict) and set(value) == keys, 'receipt-schema') + require(value['kind'] == 'native-codex-core-coding' + and value['model'] == 'qwen3-0.6b-v1' and value['phase'] in PHASES + and value['diagnostic'] in DIAGNOSTICS and all(type(value[key]) is bool for key in booleans), + 'receipt-values') + require(value['full_native_prompt_sha256'] == PROMPT_SHA256 + and value['before_sha256'] == hashlib.sha256(ORIGINAL).hexdigest() + and (value['after_sha256'] is None or isinstance(value['after_sha256'], str) + and re.fullmatch('[0-9a-f]{64}', value['after_sha256'])) + and value['private_peer_execution_claimed'] is False + and value['general_coding_quality_claimed'] is False, 'receipt-scope') + for key in ('accepted_commands', 'declined_commands'): + require(type(value[key]) is int and 0 <= value[key] <= 16, 'receipt-count') + if value['version'] >= 2: + denials = value['approval_denials'] + require(isinstance(denials, dict) and set(denials) == set(APPROVAL_DENIALS) + and all(type(count) is int and 0 <= count <= 16 for count in denials.values()) + and sum(denials.values()) == value['declined_commands'], 'receipt-approval-denials') + require(value['runtime_exit'] is None or type(value['runtime_exit']) is int + and -255 <= value['runtime_exit'] <= 255, 'receipt-exit') + counters = value['responses'] + require(counters is None or isinstance(counters, dict) + and set(counters) == {'submitted', 'completed', 'incomplete', 'cleanup_confirmed'} + and all(type(v) is int and 0 <= v <= 32 for v in counters.values()), 'receipt-responses') + if value['version'] == 3: + check_native_diagnostics(value) + if value['success']: + require(all(value[key] for key in ('native_turn_completed', 'read', 'edit', 'test', + 'independent_test_passed', 'thread_unsubscribed')) + and value['phase'] == 'complete' and value['runtime_exit'] == 0 + and not value['unexpected_command'] and not value['forced_stop'] and value['diagnostic'] is None + and value['after_sha256'] not in (None, value['before_sha256']) + and counters is not None and counters['completed'] >= 4 and counters['incomplete'] == 0 + and counters['submitted'] == counters['completed'] == counters['cleanup_confirmed'], + 'receipt-success-unproven') + if value['version'] == 3: + require(1 <= value['turns_started'] == value['turns_completed'] <= 2 + and value['response_diagnostics']['truncated'] is False + and value['item_types']['commandExecution'] >= 3, 'receipt-native-task-unproven') + return value + + +def interrupted(_signum, _frame): + raise KeyboardInterrupt + + +def run(args): + require(os.getuid() != 0, 'root-refused') + binary = verified_file(args.app_server, args.app_server_sha256, executable=True) + provenance = verified_build(args.build_report, binary, args.app_server_sha256) + node = verified_file(args.node, args.node_sha256, executable=True) + prompt = verified_file(args.upstream_prompt, PROMPT_SHA256) + ipc = private_socket(args.socket) + require(prompt.stat().st_size == 20903, 'native-prompt-size') + home = pwd.getpwuid(os.getuid()).pw_dir + require(Path(home).parent == Path('/home') and Path(home).name not in ('', '.', '..'), 'account-home') + output = Path(args.output) + require(output.is_absolute() and output.parent.resolve(strict=True) == output.parent + and output.parent.is_relative_to(ROOT / 'build') and not output.exists() + and not output.is_symlink(), 'new-workspace-output-required') + plan = dict(kind='native-codex-core-coding', upstream_revision=UPSTREAM, + app_server_sha256=args.app_server_sha256, node_sha256=args.node_sha256, + native_prompt_sha256=PROMPT_SHA256, model='qwen3-0.6b-v1', + model_download=False, service_started=False, private_peer_execution_claimed=False, + actions=['connect to existing same-owner private core', 'isolate synthetic arithmetic project', + 'native model-driven read/edit/test with exact per-command approval', + 'independent arithmetic tests', 'unsubscribe and stop runtime', 'remove private fixture']) + print(json.dumps(plan), flush=True) + if not args.execute: + require(not args.yes, 'execute-required') + return 0 + require(args.yes, 'explicit-confirmation-required') + output.mkdir(mode=0o700) + work = Path(tempfile.mkdtemp(prefix='private-', dir=output)) + (work / 'ipc').mkdir(mode=0o700) + (work / 'ipc/compute.sock').touch(mode=0o600) + (work / 'state').mkdir(mode=0o700) + (work / 'state/project').mkdir(mode=0o700) + (work / 'state/project/arithmetic.py').write_bytes(ORIGINAL) + (work / 'state/project/arithmetic.py').chmod(0o600) + before = snapshot() + result = dict(version=1, kind=plan['kind'], success=False, plan=plan, before=before, + runtime_provenance=provenance, + source_sha256={str(path.relative_to(ROOT)): digest(path) for path in + [*(ROOT / 'src' / name for name in SOURCES), ROOT / 'scripts/smoke_native_coding.cjs', + ROOT / 'scripts/native_coding_fixture.py', Path(__file__).resolve()]}) + process = None + previous = {number: signal.getsignal(number) for number in (signal.SIGINT, signal.SIGTERM, signal.SIGHUP)} + for number in previous: + signal.signal(number, interrupted) + try: + process = subprocess.Popen(command(binary, node, ipc, prompt, work, home), + stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, start_new_session=True) + result['exit_code'] = process.wait(timeout=2500) + result['observed'] = closed_receipt(work / 'state/receipt.json') + result['success'] = result['exit_code'] == 0 and result['observed']['success'] + except (OSError, ValueError, subprocess.TimeoutExpired, KeyboardInterrupt): + result['failure'] = 'native-coding-incomplete' + finally: + for number in previous: + signal.signal(number, signal.SIG_IGN) + if process is not None and process.poll() is None: + os.killpg(process.pid, signal.SIGTERM) + try: + process.wait(timeout=10) + except subprocess.TimeoutExpired: + os.killpg(process.pid, signal.SIGKILL) + process.wait(timeout=5) + result['success'] = False + joined = process is None or process.poll() is not None + if joined: + shutil.rmtree(work) + result['after'] = snapshot() + result['cleanup'] = dict(process_joined=joined, private_state_removed=not work.exists(), + host_network_unchanged=before == result['after']) + result['success'] &= all(result['cleanup'].values()) + with (output / 'report.json').open('x') as target: + json.dump(result, target, indent=2) + target.write('\n') + (output / 'report.json').chmod(0o600) + for number, handler in previous.items(): + signal.signal(number, handler) + print(json.dumps({'success': result['success'], 'cleanup': result['cleanup']}), flush=True) + return 0 if result['success'] else 1 + + +def main(): + if len(sys.argv) == 3 and sys.argv[1] == '--inside': + return inside(sys.argv[2]) + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument('--app-server', required=True) + parser.add_argument('--app-server-sha256', required=True) + parser.add_argument('--build-report', required=True) + parser.add_argument('--node', required=True) + parser.add_argument('--node-sha256', required=True) + parser.add_argument('--upstream-prompt', required=True) + parser.add_argument('--socket', required=True) + parser.add_argument('--output', required=True) + parser.add_argument('--execute', action='store_true') + parser.add_argument('--yes', action='store_true') + return run(parser.parse_args()) + + +if __name__ == '__main__': + try: + raise SystemExit(main()) + except (OSError, ValueError): + print('native-coding-preflight-failed', file=sys.stderr) + raise SystemExit(1) from None diff --git a/src/app-server.cjs b/src/app-server.cjs index 7321750..ce6b77a 100644 --- a/src/app-server.cjs +++ b/src/app-server.cjs @@ -10,9 +10,17 @@ const MAX_LINE = 1024 * 1024; class AppServer extends EventEmitter { #read; #write; #buffer = Buffer.alloc(0); #next = 0; #pending = new Map(); #closed = false; #ready = false; #initializing = false; #timeout; - constructor(readable, writable, {timeoutMs = 15000} = {}) { + #models; #writableRoot; #commandApproval; #approvals = new Map(); + constructor(readable, writable, {timeoutMs = 15000, allowedModels = [], writableRoot = null, + commandApproval = null} = {}) { super(); if (!Number.isInteger(timeoutMs) || timeoutMs < 1 || timeoutMs > 60000) throw Error('rpc_timeout'); + if (!Array.isArray(allowedModels) || allowedModels.length > 8 || + !allowedModels.every(model => typeof model === 'string' && /^[a-z0-9][a-z0-9._-]{0,95}$/.test(model)) || + commandApproval !== null && typeof commandApproval !== 'function' || + writableRoot !== null && (typeof writableRoot !== 'string' || !writableRoot.startsWith('/') || + writableRoot.includes('\0') || typeof commandApproval !== 'function')) throw Error('client_scope'); + this.#models = new Set(allowedModels); this.#writableRoot = writableRoot; this.#commandApproval = commandApproval; this.#read = readable; this.#write = writable; this.#timeout = timeoutMs; readable.on('data', chunk => this.#data(chunk)); readable.on('end', () => this.close()); @@ -63,9 +71,23 @@ class AppServer extends EventEmitter { if (typeof value.method === 'string') { if (Object.hasOwn(value, 'id')) { if (!(typeof value.id === 'string' && value.id.length <= 256) && !Number.isSafeInteger(value.id)) throw Error('rpc_id'); - // No automatic approval, tools, authentication or peer dispatch. A later - // interactive approval adapter must explicitly replace this fail-closed behavior. - if (['item/commandExecution/requestApproval', 'item/fileChange/requestApproval'].includes(value.method)) { + // Defaults remain read-only/decline. Only an explicitly supplied caller + // policy may approve one command; never approve sessions or policy changes. + if (value.method === 'item/commandExecution/requestApproval' && this.#commandApproval && this.#ready) { + if (this.#approvals.size >= 8 || this.#approvals.has(value.id)) throw Error('approval_bound'); + const finish = accepted => { + const timer = this.#approvals.get(value.id); + if (timer === undefined) return; + clearTimeout(timer); this.#approvals.delete(value.id); + if (!this.#closed) { + try { this.#send({id: value.id, result: {decision: accepted === true ? 'accept' : 'decline'}}); } + catch { this.close(); } + } + }; + this.#approvals.set(value.id, setTimeout(() => finish(false), 30000)); + Promise.resolve().then(() => this.#commandApproval(value.params)).then(finish, () => finish(false)) + .catch(() => this.close()); + } else if (['item/commandExecution/requestApproval', 'item/fileChange/requestApproval'].includes(value.method)) { this.#send({id: value.id, result: {decision: 'decline'}}); } else this.#send({id: value.id, error: {code: -32601, message: 'Client method not enabled'}}); } else this.emit('notification', {method: value.method, params: value.params}); @@ -87,10 +109,11 @@ class AppServer extends EventEmitter { return result; } async startThread({model, cwd}) { - if (!this.#ready || typeof model !== 'string' || !/^volparossa-[a-z0-9_-]{1,64}$/.test(model) || + if (!this.#ready || typeof model !== 'string' || + !(/^volparossa-[a-z0-9_-]{1,64}$/.test(model) || this.#models.has(model)) || typeof cwd !== 'string' || !cwd.startsWith('/') || cwd.includes('\0')) throw Error('thread_scope'); return this.request('thread/start', {model, modelProvider: 'volparossa', cwd, - sandbox: 'read-only', approvalPolicy: 'untrusted', ephemeral: true}); + sandbox: this.#writableRoot === cwd ? 'workspace-write' : 'read-only', approvalPolicy: 'untrusted', ephemeral: true}); } async startTurn(threadId, text) { if (!this.#ready || typeof threadId !== 'string' || !threadId || threadId.length > 256 || @@ -106,6 +129,8 @@ class AppServer extends EventEmitter { close() { if (this.#closed) return; this.#closed = true; this.#buffer = Buffer.alloc(0); + for (const timer of this.#approvals.values()) clearTimeout(timer); + this.#approvals.clear(); for (const {reject, timer} of this.#pending.values()) { clearTimeout(timer); reject(Error('rpc_closed')); } this.#pending.clear(); this.#read.destroy(); this.#write.destroy(); this.emit('closed'); } diff --git a/src/native-coding-fixture.cjs b/src/native-coding-fixture.cjs new file mode 100644 index 0000000..afcc9da --- /dev/null +++ b/src/native-coding-fixture.cjs @@ -0,0 +1,202 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Explicit synthetic-workspace policy, not a general editor approval policy. +'use strict'; +const {createHash} = require('node:crypto'); +const MODEL = 'qwen3-0.6b-v1'; +const PROJECT = '/opt/work/project'; +const PROMPT_SHA256 = 'ac8ae107a0d72fe3476b430afb161ea4e67da2e446d778aefc44828160559807'; + +function modelCatalog(instructions) { + if (Buffer.byteLength(instructions) !== 20903 || createHash('sha256').update(instructions).digest('hex') !== PROMPT_SHA256) { + throw Error('native_prompt_mismatch'); + } + return {models: [{slug: MODEL, display_name: 'VOLPAROSSA Qwen3 0.6B', description: null, + supported_reasoning_levels: [], default_reasoning_level: null, shell_type: 'unified_exec', + visibility: 'list', supported_in_api: true, priority: 1, upgrade: null, + model_messages: {instructions_template: instructions}, include_apps_usage_instructions: false, + supports_reasoning_summary_parameter: false, default_reasoning_summary: 'none', + support_verbosity: false, default_verbosity: null, apply_patch_tool_type: null, + truncation_policy: {mode: 'bytes', limit: 65536}, context_window: 32768, max_context_window: 32768, + effective_context_window_percent: 95, experimental_supported_tools: [], input_modalities: ['text'], + tool_mode: 'direct', use_responses_lite: false}]}; +} + +function runtimeSettings(baseUrl) { + if (!/^http:\/\/127\.0\.0\.1:[1-9][0-9]{0,4}\/v1$/.test(baseUrl)) throw Error('provider_scope'); + return [`model="${MODEL}"`, 'model_provider="volparossa"', 'model_catalog_json="/opt/catalog.json"', + 'model_providers.volparossa.name="VOLPAROSSA"', `model_providers.volparossa.base_url="${baseUrl}"`, + 'model_providers.volparossa.env_key="VOLPAROSSA_PROVIDER_TOKEN"', + 'model_providers.volparossa.wire_api="responses"', 'model_providers.volparossa.requires_openai_auth=false', + 'model_providers.volparossa.supports_websockets=false', 'model_providers.volparossa.request_max_retries=0', + 'model_providers.volparossa.stream_max_retries=0', 'model_providers.volparossa.stream_idle_timeout_ms=610000', + 'model_reasoning_summary="none"', 'check_for_update_on_startup=false', 'analytics.enabled=false', + 'feedback.enabled=false', 'web_search="disabled"', 'mcp_servers={}', 'tools.update_plan.enabled=false', + 'tools.experimental_request_user_input.enabled=false', 'features.view_image=false', + 'features.code_mode=false', 'features.code_mode_host=false', 'features.multi_agent=false', + 'features.multi_agent_v2=false', 'features.apps=false', 'features.plugins=false', + 'features.tool_suggest=false', 'features.tool_search=false', 'features.shell_snapshot=false', + 'features.exec_permission_approvals=false', 'features.request_permissions_tool=false', + 'features.unified_exec_tty=false', 'shell_environment_policy.exclude=["VOLPAROSSA_PROVIDER_TOKEN"]']; +} + +// Decode only the native shlex-joined display; this function never executes it. +function shellWords(text) { + if (typeof text !== 'string' || text.length > 4096 || /[\0\r\n]/.test(text)) return null; + const words = []; let value = '', quote = null, present = false; + for (let at = 0; at < text.length; at++) { + const c = text[at]; + if (quote === "'") { if (c === "'") quote = null; else value += c; } + else if (quote === '"') { + if (c === '"') quote = null; + else if (c === '\\') { if (++at === text.length) return null; value += text[at]; } + else value += c; + } else if (/\s/.test(c)) { + if (present) { words.push(value); value = ''; present = false; } + } else if (c === "'" || c === '"') { quote = c; present = true; } + else if (c === '\\') { if (++at === text.length) return null; value += text[at]; present = true; } + else { value += c; present = true; } + } + if (quote) return null; + if (present) words.push(value); + return words; +} + +function commandKind(display) { + const words = shellWords(display); + if (!words) return null; + let command = display; + if (words.length === 3 && ['/bin/bash', '/usr/bin/bash', '/bin/sh', '/usr/bin/sh'].includes(words[0]) && + ['-c', '-lc'].includes(words[1])) command = words[2]; + if (command === 'python3 -B /opt/fixture.py read') return 'read'; + if (command === 'python3 -B /opt/fixture.py test') return 'test'; + if (/^python3 -B \/opt\/fixture\.py edit '[ab0-9 ()+*/-]{1,80}'$/.test(command)) return 'edit'; + return null; +} + +const APPROVAL_DENIALS = Object.freeze(['lineage', 'kind', 'item', 'cwd', 'command', + 'network', 'permissions', 'network_policy', 'order', 'budget']); + +// Closed reasons only: never return commands, paths, IDs or permission contents. +function approvalDenial(params, threadId, turnId) { + if (!threadId || !turnId || !params || params.threadId !== threadId || params.turnId !== turnId) return 'lineage'; + if (params.kind !== 'command') return 'kind'; + if (typeof params.itemId !== 'string' || params.itemId.length > 256) return 'item'; + if (params.cwd !== PROJECT) return 'cwd'; + if (!commandKind(params.command)) return 'command'; + if (params.networkApprovalContext) return 'network'; + if (params.additionalPermissions) return 'permissions'; + if (params.proposedNetworkPolicyAmendments) return 'network_policy'; + // The pinned native runtime also proposes an execpolicy rule for ordinary + // commands. This is not an extra permission request: AppServer returns only + // one-shot "accept"/"decline", never acceptWithExecpolicyAmendment or session + // approval. Ignore the proposal; it neither authorizes nor changes a command. + return null; +} + +function authorize(params, threadId, turnId) { + return approvalDenial(params, threadId, turnId) === null; +} + +const TASK = `Fix the add(a,b) function in the synthetic fixture, then verify it. Do not guess its current source. +Use native exec_command, workdir ${PROJECT}, shell /bin/bash, login false, tty false, max_output_tokens 1024. +Only these command forms are authorized, one at a time: +1. python3 -B /opt/fixture.py read +2. python3 -B /opt/fixture.py edit 'EXPRESSION' (replace EXPRESSION with your arithmetic expression in a and b) +3. python3 -B /opt/fixture.py test +The edit helper writes exactly your proposed expression, not a predetermined repair. It accepts only arithmetic. +Read first, inspect the returned source, make the minimal edit, and run the actual tests. Finish only after tests pass. +Do not execute other commands, request escalation, alter tests, use network, or invent tool results.`; +const CONTINUATION = 'The original task is not complete according to the observed read, edit and test results. Continue the original task in this same workspace, within the same permissions. Do not claim completion until the required work and actual verification are complete.'; +const ITEM_TYPES = Object.freeze(['commandExecution', 'agentMessage', 'userMessage', 'reasoning', 'other']); + +// One task may require a second normal native turn. Neither turn completion nor +// prose is task evidence. This controller never supplies a repair or executes tools. +class NativeTaskController { + constructor(client, threadId, report) { + this.client = client; this.threadId = threadId; this.report = report; + this.turnId = null; this.terminal = null; this.pending = false; + this.stopped = false; this.nativeItems = 0; this.complete = () => {}; + this.notify = value => this.notification(value); + this.closed = () => { this.stopped = true; this.complete(); }; + client.on('notification', this.notify); client.on('closed', this.closed); + } + dispose() { + this.client.off('notification', this.notify); this.client.off('closed', this.closed); + } + async stop() { + if (this.stopped) return; + this.stopped = true; this.complete(); + if (this.pending && this.turnId) { + try { await this.client.interrupt(this.threadId, this.turnId); } catch {} + } + } + approve(params) { + const kind = commandKind(params?.command), report = this.report; + const inOrder = kind === 'read' || kind === 'edit' && report.read || kind === 'test' && report.read && report.edit; + const denial = approvalDenial(params, this.threadId, this.pending && !this.stopped && !this.terminal ? this.turnId : null) ?? + (!inOrder ? 'order' : report.accepted_commands >= 6 ? 'budget' : null); + if (denial === null) report.accepted_commands++; + else { report.declined_commands++; report.approval_denials[denial]++; } + if (report.declined_commands > 2) void this.stop(); + return denial === null; + } + notification({method, params}) { + if (!this.pending || this.stopped || this.terminal || params?.threadId !== this.threadId) return; + if (method === 'turn/started') { + const id = params.turn?.id; + if (typeof id !== 'string' || !id || id.length > 256 || this.turnId && this.turnId !== id) { + void this.stop(); return; + } + this.turnId = id; + } + if (method === 'item/completed' && params.turnId === this.turnId) { + const report = this.report, item = params.item; + const type = ITEM_TYPES.includes(item?.type) ? item.type : 'other'; + if (Object.values(report.item_types).reduce((a, b) => a + b, 0) >= 64) { void this.stop(); return; } + report.item_types[type]++; + if (item?.type !== 'commandExecution') return; + if (++this.nativeItems > 8) { void this.stop(); return; } + const kind = commandKind(item.command); + if (!kind || item.cwd !== PROJECT || kind === 'edit' && !report.read || kind === 'test' && !report.edit) { + report.unexpected_command = true; void this.stop(); return; + } + if (kind === 'edit') report.test = false; + if (item.status === 'completed' && item.exitCode === 0) report[kind] = true; + } + if (method === 'turn/completed' && this.turnId && params.turn?.id === this.turnId) { + this.terminal = params.turn; this.complete(); + } + } + async run() { + try { + for (const input of [TASK, CONTINUATION]) { + if (this.stopped) throw Error('native_task_stopped'); + this.turnId = null; this.terminal = null; this.pending = true; + this.report.native_turn_completed = false; + const finished = new Promise(resolve => { this.complete = resolve; }); + this.report.turns_started++; + const admitted = await this.client.startTurn(this.threadId, input); + const id = admitted?.turn?.id; + if (typeof id !== 'string' || !id || id.length > 256 || this.turnId && this.turnId !== id) { + throw Error('native_turn_mismatch'); + } + this.turnId = id; + await finished; + this.pending = false; + if (this.stopped || this.terminal?.status !== 'completed' || this.terminal.error) { + throw Error('native_turn_incomplete'); + } + this.report.turns_completed++; + this.report.native_turn_completed = true; + if (this.report.read && this.report.edit && this.report.test && !this.report.unexpected_command) return; + if (this.report.accepted_commands >= 6 || this.report.unexpected_command) break; + } + throw Error('native_task_incomplete'); + } catch (error) { + await this.stop(); + throw error; + } finally { this.pending = false; } + } +} +module.exports = {MODEL, PROJECT, PROMPT_SHA256, modelCatalog, runtimeSettings, commandKind, + APPROVAL_DENIALS, approvalDenial, authorize, TASK, CONTINUATION, ITEM_TYPES, NativeTaskController}; diff --git a/src/private-conversation.cjs b/src/private-conversation.cjs new file mode 100644 index 0000000..5e55d3f --- /dev/null +++ b/src/private-conversation.cjs @@ -0,0 +1,202 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Additive conversation protocol. Reuse the unchanged Q&A framing/owner/cancel transport. +'use strict'; + +const { PrivateCompute } = require('./private-compute.cjs'); + +function fail(code) { + const error = new Error(`private_compute_${code}`); + error.code = code; + if (code === 'cancelled') error.name = 'AbortError'; + return error; +} +function check(value, code = 'invalid_conversation') { if (!value) throw fail(code); } +function object(value) { return value !== null && typeof value === 'object' && !Array.isArray(value); } +function keys(value, required, optional = []) { + check(object(value) && required.every(key => Object.hasOwn(value, key)) && + Object.keys(value).every(key => required.includes(key) || optional.includes(key))); +} +function text(value, maximum, nonempty = true) { + check(typeof value === 'string' && Buffer.byteLength(value) <= maximum && + Buffer.from(value).toString() === value && !value.includes('\0') && (!nonempty || value.trim())); +} +function identifier(value) { check(typeof value === 'string' && /^[A-Za-z0-9_.-]{1,64}$/.test(value)); } +function toolKey(value) { + identifier(value.name); + check(value.namespace == null || typeof value.namespace === 'string'); + if (value.namespace != null) identifier(value.namespace); + return JSON.stringify([value.namespace ?? null, value.name]); +} +function payload(value) { check(object(value) && Buffer.byteLength(JSON.stringify(value)) <= 4096); } + +const PROFILES = Object.freeze({ + 'smollm2-135m-v1': [192, 64, 1024], + 'smollm2-360m-v1': [1024, 256, 4096], + 'smollm2-1.7b-v1': [1024, 256, 4096], + 'qwen3-0.6b-v1': [12288, 1024, 4096], +}); +function expectedLimits(model) { + check(Object.hasOwn(PROFILES, model), 'incompatible_capabilities'); + const [prompt, output, bytes] = PROFILES[model]; + const qwen = model === 'qwen3-0.6b-v1'; + return { version: 1, visibility: 'private_local', model_profile: model, + max_input_bytes: 24576, max_history_items: 32, max_tools: 8, + max_instructions_bytes: 4096, max_message_bytes: 8192, max_tool_description_bytes: 2048, + max_tool_payload_bytes: 4096, max_prompt_tokens: prompt, max_new_tokens: output, + model_context_tokens: 8192, max_output_bytes: bytes, conversation_template: 'smollm2-json-turn-v1', + native_tool_template: false, local_only: true, tool_execution: false, network_access: false, + public_cache: false, training: false, cloud_fallback: false, model_tool_use_proven: false, + arbitrary_json_schema_validation: false, + ...(qwen ? { max_input_bytes: 262144, max_instructions_bytes: 65536, max_history_items: 128, + max_tools: 32, max_message_bytes: 65536, max_tool_description_bytes: 8192, + model_context_tokens: 32768, conversation_template: 'qwen3-tools-nonthinking-v1', native_tool_template: true } : {}) }; +} +function requestLimit(model) { return model === 'qwen3-0.6b-v1' ? 524288 : 32768; } +function equalLimits(value, expected, optional = []) { + keys(value, Object.keys(expected), optional); + check(Object.entries(expected).every(([key, item]) => value[key] === item), 'incompatible_capabilities'); +} +function capabilities(value) { + check(object(value)); + equalLimits(value, { ...expectedLimits(value.model_profile), execution_slots: 1, + max_request_bytes: requestLimit(value.model_profile), max_response_bytes: 65536 }, ['max_seconds', 'quarantined']); + check(Number.isInteger(value.max_seconds) && value.max_seconds >= 1 && value.max_seconds <= 600 && + typeof value.quarantined === 'boolean', 'incompatible_capabilities'); + return Object.freeze(value); +} + +function validateConversation(value, limits = expectedLimits('smollm2-360m-v1')) { + keys(value, ['version', 'visibility', 'instructions', 'history', 'tools']); + check(value.version === 1 && value.visibility === 'private_local'); + text(value.instructions, limits.max_instructions_bytes); + check(Array.isArray(value.history) && value.history.length >= 1 && value.history.length <= limits.max_history_items && + Array.isArray(value.tools) && value.tools.length <= limits.max_tools); + const tools = new Map(), seen = new Map(), open = new Set(); + for (const tool of value.tools) { + check(tool.type === 'function' || tool.type === 'custom'); + keys(tool, ['type', 'name', 'description', ...(tool.type === 'function' ? ['parameters'] : [])], ['namespace']); + const key = toolKey(tool); + check(!tools.has(key)); + tools.set(key, tool.type); + text(tool.description, limits.max_tool_description_bytes); + if (tool.type === 'function') { payload(tool.parameters); check(tool.parameters.type === 'object'); } + } + for (const item of value.history) { + check(object(item)); + if (item.type === 'message') { + keys(item, ['type', 'role', 'text']); + const roles = limits.model_profile === 'qwen3-0.6b-v1' ? ['user', 'assistant', 'system', 'developer'] : ['user', 'assistant']; + check(!open.size && roles.includes(item.role)); + text(item.text, limits.max_message_bytes); + } else if (item.type === 'tool_result') { + keys(item, ['type', 'call_id', 'output']); + identifier(item.call_id); + check(open.delete(item.call_id)); + text(item.output, limits.max_message_bytes, false); + } else { + validateCall(item, tools, seen); + seen.set(item.call_id, item); + open.add(item.call_id); + } + } + const last = value.history.at(-1); + check(!open.size && (last.type === 'tool_result' || last.type === 'message' && last.role === 'user')); + check(Buffer.byteLength(JSON.stringify(value)) <= limits.max_input_bytes, 'request_bound'); + return { tools, seen }; +} +function validateCall(item, tools, seen) { + const custom = item.type === 'custom_tool_call'; + check(custom || item.type === 'function_call'); + keys(item, ['type', 'call_id', 'name', custom ? 'input' : 'arguments'], ['namespace']); + identifier(item.call_id); + check(!seen.has(item.call_id) && tools.get(toolKey(item)) === (custom ? 'custom' : 'function')); + if (custom) text(item.input, 4096); else payload(item.arguments); +} +function validateResult(value, caps, input) { + keys(value, ['version', 'operation', 'model_profile', 'execution_complete', 'turn_complete', 'output', + 'prompt_tokens', 'generated_tokens', 'limits', 'local_only', 'private_data_supported', 'tool_execution', + 'distributed_execution_claimed', 'private_training_claimed', 'model_answer_correctness_proven', 'cleanup']); + keys(value.cleanup, ['complete', 'retained_input', 'retained_report']); + check(value.cleanup.complete === true && value.cleanup.retained_input === false && + value.cleanup.retained_report === false, 'cleanup_unconfirmed'); + check(value.version === 1 && value.operation === 'compute_private_conversation' && + value.model_profile === caps.model_profile && value.execution_complete === true && value.local_only === true && + value.private_data_supported === true && value.tool_execution === false && + value.distributed_execution_claimed === false && value.private_training_claimed === false && + value.model_answer_correctness_proven === false); + equalLimits(value.limits, expectedLimits(caps.model_profile)); + check(Number.isInteger(value.prompt_tokens) && value.prompt_tokens >= 1 && value.prompt_tokens <= caps.max_prompt_tokens && + Number.isInteger(value.generated_tokens) && value.generated_tokens >= 1 && value.generated_tokens <= caps.max_new_tokens); + const output = value.output; + check(object(output) && value.turn_complete === (output.type !== 'incomplete')); + if (output.type === 'incomplete') { + keys(output, ['type', 'reason']); + check(['token_limit', 'wire_truncated', 'invalid_output'].includes(output.reason)); + } else if (output.type === 'assistant') { + keys(output, ['type', 'text']); + text(output.text, caps.max_output_bytes); + } else { + const { tools, seen } = validateConversation(input, caps); + check(Object.hasOwn(output, 'namespace')); + validateCall(output, tools, seen); + } + return value; +} + +class PrivateConversation extends PrivateCompute { + // These hooks reuse only the already-tested transport. Q&A callers and bytes + // remain unchanged; this instance cannot silently use the old handshake. + _send(id, operation) { + // The separate conversation family may advertise a larger request envelope; + // the legacy Q&A class and its 32KiB framing remain byte-for-byte unchanged. + if (operation.type === 'capabilities') operation = { type: 'conversation_capabilities' }; + try { + const body = Buffer.from(JSON.stringify({ version: 1, id, operation })); + check(body.length > 0 && body.length <= (this.caps?.max_request_bytes ?? 32768), 'request_bound'); + check(this.socket && !this.socket.destroyed, 'disconnected'); + const header = Buffer.alloc(4); header.writeUInt32BE(body.length); + this.socket.write(Buffer.concat([header, body])); + } catch (error) { this._fail(error.message?.startsWith('private_compute_') ? error : fail('socket_error')); } + } + ask() { return Promise.reject(fail('conversation_required')); } + async submit(conversation, { signal } = {}) { + check(this.state === 'open', 'not_connected'); + validateConversation(conversation, this.caps); + check(signal === undefined || signal instanceof AbortSignal, 'invalid_signal'); + check(!this.caps.quarantined, 'cleanup_unconfirmed'); + check(!this.pending, 'busy'); + if (signal?.aborted) throw fail('cancelled'); + const id = this._id(); + // Snapshot before any asynchronous wait; caller mutations cannot change + // correlation/tool validation after the exact submitted frame has left. + const input = JSON.parse(JSON.stringify(conversation)); + return new Promise((resolve, reject) => { + const abort = () => this._cancel(); + this.pending = { id, input, resolve, reject, signal, abort, admitted: false, cancelled: false, + timer: setTimeout(() => this._cancel(), this.caps.max_seconds * 1000 + 5000) }; + signal?.addEventListener('abort', abort, { once: true }); + this._send(id, { type: 'submit_conversation', conversation: input }); + if (signal?.aborted) abort(); + }); + } + _response(message) { + check(object(message) && message.version === 1 && message.event !== 'capabilities', 'invalid_response'); + if (message.event === 'conversation_capabilities') { + keys(message, ['version', 'id', 'event', 'capabilities']); + check(this.state === 'connecting' && message.id === this.handshake?.id, 'invalid_response'); + this.caps = capabilities(message.capabilities); + this.state = 'open'; + clearTimeout(this.handshake.timer); + this.handshake.resolve(this.caps); + this.handshake = null; + } else if (message.event === 'result') { + keys(message, ['version', 'id', 'event', 'result']); + check(this.pending?.admitted && message.id === this.pending.id, 'invalid_response'); + const result = validateResult(message.result, this.caps, this.pending.input); + this._settle(this.pending.cancelled ? fail('cancelled') : null, result); + } else super._response(message); + } +} + +module.exports = { PrivateConversation, validateConversation, validateResult, expectedLimits, requestLimit, capabilities, + check, keys, text, identifier, object, fail }; diff --git a/src/responses-provider.cjs b/src/responses-provider.cjs new file mode 100644 index 0000000..123ed21 --- /dev/null +++ b/src/responses-provider.cjs @@ -0,0 +1,313 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Explicit local Responses subset. No model, cloud client, scheduler or tool execution. +'use strict'; + +const http = require('node:http'); +const { randomBytes, timingSafeEqual } = require('node:crypto'); +const { TextDecoder } = require('node:util'); +const { PrivateConversation, validateConversation, expectedLimits, check, keys, text, + identifier, object, fail } = require('./private-conversation.cjs'); + +const HTTP_BYTES = 524288; + +// JSON.parse alone silently discards duplicate keys (including inside function +// argument strings). Preserve one unambiguous request without repair or truncation. +function parseJson(raw) { + text(raw, HTTP_BYTES, false); + let value; + try { value = JSON.parse(raw); } catch { throw fail('invalid_json'); } + const tokens = raw.match(/"(?:[^"\\]|\\.)*"|[{}\[\],:]|true|false|null|-?\d+(?:\.\d+)?(?:[eE][+-]?\d+)?/g) ?? []; + for (const token of tokens) { + if (/^-?\d/.test(token)) { + const number = Number(token); + check(Number.isFinite(number) && (!Number.isInteger(number) || Number.isSafeInteger(number)), 'json_number_bound'); + } + } + let at = 0, nodes = 0; + function visit(depth) { + check(depth <= 32 && ++nodes <= 8192, 'json_bound'); + const token = tokens[at++]; + if (token === '{') { + const seen = new Set(); + if (tokens[at] === '}') { at++; return; } + do { + const key = JSON.parse(tokens[at++]); + check(!seen.has(key), 'duplicate_json_key'); + seen.add(key); + check(tokens[at++] === ':', 'invalid_json'); + visit(depth + 1); + } while (tokens[at++] === ','); + } else if (token === '[') { + if (tokens[at] === ']') { at++; return; } + do { visit(depth + 1); } while (tokens[at++] === ','); + } + } + visit(0); + check(at === tokens.length, 'invalid_json'); + return value; +} + +function content(value, role) { + if (typeof value === 'string') return value; + check(Array.isArray(value) && value.length >= 1 && value.length <= 128, 'unsupported_content'); + return value.map(part => { + keys(part, ['type', 'text'], ['annotations', 'logprobs']); + check(part.type === (role === 'assistant' ? 'output_text' : 'input_text') && + (part.annotations === undefined || Array.isArray(part.annotations) && !part.annotations.length) && + (part.logprobs === undefined || Array.isArray(part.logprobs) && !part.logprobs.length), 'unsupported_content'); + check(typeof part.text === 'string', 'unsupported_content'); + return part.text; + }).join('\n\n'); +} +function neutral(value, key, allowed) { + check(value[key] === undefined || allowed.some(item => JSON.stringify(value[key]) === JSON.stringify(item)), 'unsupported_feature'); +} +function normalizeTools(raw, caps) { + check(Array.isArray(raw) && raw.length <= caps.max_tools, 'tool_bound'); + const result = []; + function add(tool, namespace = null, description = '') { + check(object(tool) && ['function', 'custom'].includes(tool.type), 'unsupported_tool'); + keys(tool, ['type', 'name', 'description', ...(tool.type === 'function' ? ['parameters'] : [])], + ['namespace', 'strict', 'format', 'defer_loading']); + check(namespace === null || tool.namespace === undefined, 'namespace_conflict'); + neutral(tool, 'strict', [false]); + neutral(tool, 'defer_loading', [false]); + if (tool.type === 'custom') neutral(tool, 'format', [{ type: 'text' }]); + else check(tool.format === undefined, 'unsupported_feature'); + text(tool.description, caps.max_tool_description_bytes); + const converted = { type: tool.type, name: tool.name, namespace: namespace ?? tool.namespace ?? null, + description: description ? `Namespace ${namespace}: ${description}\n\n${tool.description}` : tool.description }; + if (tool.type === 'function') converted.parameters = tool.parameters; + result.push(converted); + check(result.length <= caps.max_tools, 'tool_bound'); + } + for (const tool of raw) { + if (tool?.type !== 'namespace') { add(tool); continue; } + keys(tool, ['type', 'name', 'description', 'tools']); + identifier(tool.name); + text(tool.description, caps.max_tool_description_bytes, false); + check(Array.isArray(tool.tools) && tool.tools.length >= 1 && tool.tools.length <= caps.max_tools, 'tool_bound'); + for (const member of tool.tools) add(member, tool.name, tool.description); + } + return result; +} + +function toConversation(request, model, caps) { + keys(request, ['model', 'instructions', 'input', 'store', 'stream'], ['tools', 'tool_choice', 'parallel_tool_calls', + 'reasoning', 'include', 'text', 'stream_options', 'max_output_tokens', 'prompt_cache_key', 'client_metadata', 'service_tier']); + check(request.model === model && request.store === false && request.stream === true, 'unsupported_request'); + neutral(request, 'tool_choice', ['auto']); + neutral(request, 'parallel_tool_calls', [false, true]); // Permission for several is not a requirement to propose several. + if (request.reasoning != null) { + keys(request.reasoning, [], ['effort', 'summary', 'context']); + for (const [key, allowed] of [['effort', ['none']], ['summary', ['none']], ['context', ['current_turn']]]) { + check(request.reasoning[key] === undefined || allowed.includes(request.reasoning[key]), 'unsupported_reasoning'); + } + } + // The pinned Codex builder always asks for this optional field, even on a + // non-reasoning provider. There is no reasoning item/encrypted state to return. + neutral(request, 'include', [[], ['reasoning.encrypted_content']]); + neutral(request, 'text', [null, {}]); + neutral(request, 'stream_options', [null, { include_usage: true }]); + neutral(request, 'service_tier', [null, 'default']); + // Explicitly accepted transport hints are discarded, never passed to a model, + // cached, persisted, logged or used as authority. No cache hit is claimed. + if (request.prompt_cache_key != null) text(request.prompt_cache_key, 512); + if (request.client_metadata != null) { + check(object(request.client_metadata) && Object.keys(request.client_metadata).length <= 32, 'metadata_bound'); + for (const [key, value] of Object.entries(request.client_metadata)) { text(key, 128); text(value, 16384, false); } + check(Buffer.byteLength(JSON.stringify(request.client_metadata)) <= 32768, 'metadata_bound'); + } + if (request.max_output_tokens !== undefined) check(request.max_output_tokens === caps.max_new_tokens, 'unsupported_output_budget'); + const tools = normalizeTools(request.tools ?? [], caps); + const raw = typeof request.input === 'string' ? [{ type: 'message', role: 'user', content: request.input }] : request.input; + check(Array.isArray(raw) && raw.length >= 1 && raw.length <= caps.max_history_items, 'history_bound'); + const history = [], calls = new Map(); + for (const item of raw) { + check(object(item), 'unsupported_history'); + if (item.id !== undefined) text(item.id, 128); + if (item.type === 'message' || item.type === undefined && item.role !== undefined) { + keys(item, ['role', 'content'], ['type', 'id', 'status']); + neutral(item, 'status', ['completed']); + const roles = model === 'qwen3-0.6b-v1' ? ['user', 'assistant', 'system', 'developer'] : ['user', 'assistant']; + check(roles.includes(item.role), 'unsupported_role'); + history.push({ type: 'message', role: item.role, text: content(item.content, item.role) }); + } else if (['function_call', 'custom_tool_call'].includes(item.type)) { + const custom = item.type === 'custom_tool_call'; + keys(item, ['type', 'call_id', 'name', custom ? 'input' : 'arguments'], ['id', 'namespace', 'status']); + neutral(item, 'status', ['completed']); + check(!calls.has(item.call_id), 'duplicate_call'); + const call = { type: item.type, call_id: item.call_id, name: item.name, namespace: item.namespace ?? null, + ...(custom ? { input: item.input } : { arguments: parseJson(item.arguments) }) }; + calls.set(item.call_id, call); + history.push(call); + } else if (['function_call_output', 'custom_tool_call_output'].includes(item.type)) { + keys(item, ['type', 'call_id', 'output'], ['id', 'name', 'namespace']); + const call = calls.get(item.call_id); + check(call && item.type === (call.type === 'function_call' ? 'function_call_output' : 'custom_tool_call_output') && + (item.name === undefined || item.name === call.name) && + (item.namespace === undefined || item.namespace === call.namespace), 'tool_result_correlation'); + history.push({ type: 'tool_result', call_id: item.call_id, output: content(item.output, 'user') }); + } else throw fail('unsupported_history'); + } + const conversation = { version: 1, visibility: 'private_local', instructions: request.instructions, history, tools }; + validateConversation(conversation, caps); + return conversation; +} + +function responseEvents(result) { + // Caller receives this only after PrivateConversation validates terminal cleanup. + const id = `resp_${randomBytes(16).toString('hex')}`; + const created = Math.floor(Date.now() / 1000); + const base = { id, object: 'response', created_at: created, model: result.model_profile, store: false, + error: null, incomplete_details: null, output: [] }; + let sequence = 0; + const events = []; + const emit = (type, fields) => events.push({ type, sequence_number: sequence++, ...fields }); + emit('response.created', { response: { ...base, status: 'in_progress' } }); + const usage = { input_tokens: result.prompt_tokens, output_tokens: result.generated_tokens, + total_tokens: result.prompt_tokens + result.generated_tokens, + input_tokens_details: { cached_tokens: 0 }, output_tokens_details: { reasoning_tokens: 0 } }; + if (!result.turn_complete) { + const reason = result.output.reason === 'token_limit' ? 'max_output_tokens' : result.output.reason; + emit('response.incomplete', { response: { ...base, status: 'incomplete', + incomplete_details: { reason }, usage } }); + return events; + } + const output = result.output; + let item; + if (output.type === 'assistant') { + const itemId = `msg_${randomBytes(16).toString('hex')}`; + const part = { type: 'output_text', text: output.text, annotations: [], logprobs: [] }; + item = { id: itemId, type: 'message', role: 'assistant', status: 'completed', content: [part] }; + emit('response.output_item.added', { output_index: 0, item: { ...item, status: 'in_progress', content: [] } }); + emit('response.content_part.added', { item_id: itemId, output_index: 0, content_index: 0, + part: { ...part, text: '' } }); + emit('response.output_text.delta', { item_id: itemId, output_index: 0, content_index: 0, delta: output.text, logprobs: [] }); + emit('response.output_text.done', { item_id: itemId, output_index: 0, content_index: 0, text: output.text, logprobs: [] }); + emit('response.content_part.done', { item_id: itemId, output_index: 0, content_index: 0, part }); + } else { + const custom = output.type === 'custom_tool_call'; + item = { id: `${custom ? 'ctc' : 'fc'}_${randomBytes(16).toString('hex')}`, type: output.type, + status: 'completed', call_id: output.call_id, name: output.name, namespace: output.namespace, + ...(custom ? { input: output.input } : { arguments: JSON.stringify(output.arguments) }) }; + const field = custom ? 'input' : 'arguments'; + emit('response.output_item.added', { output_index: 0, item: { ...item, status: 'in_progress', [field]: '' } }); + const prefix = custom ? 'response.custom_tool_call_input' : 'response.function_call_arguments'; + emit(`${prefix}.delta`, { item_id: item.id, output_index: 0, delta: item[field] }); + emit(`${prefix}.done`, { item_id: item.id, output_index: 0, [field]: item[field] }); + } + emit('response.output_item.done', { output_index: 0, item }); + emit('response.completed', { response: { ...base, status: 'completed', output: [item], usage, + end_turn: output.type === 'assistant' } }); + return events; +} + +function errorReply(response, status, code) { + if (response.destroyed || response.writableEnded) return; + response.writeHead(status, { 'Content-Type': 'application/json', 'Cache-Control': 'no-store', 'Connection': 'close' }); + response.end(JSON.stringify({ error: { type: 'volparossa_provider_error', code, + message: 'Local VOLPAROSSA provider rejected or could not complete this operation.' } })); +} +async function readBody(request) { + const chunks = []; + let size = 0; + for await (const chunk of request) { + size += chunk.length; + check(size <= HTTP_BYTES, 'request_bound'); + chunks.push(chunk); + } + check(size > 0, 'invalid_request'); + try { return parseJson(new TextDecoder('utf-8', { fatal: true }).decode(Buffer.concat(chunks))); } + catch (error) { throw error.message?.startsWith('private_compute_') ? error : fail('invalid_json'); } +} + +async function startResponsesProvider({ socketPath, model, diagnostics = false }) { + expectedLimits(model); + check(typeof diagnostics === 'boolean', 'diagnostic_scope'); + // Construction checks path syntax, but creates no connection or model process. + new PrivateConversation(socketPath); + const bearerToken = randomBytes(32).toString('base64url'); + const authorization = Buffer.from(`Bearer ${bearerToken}`); + let host, active = null, closing = false; + const observed = {submitted: 0, completed: 0, incomplete: 0, cleanup_confirmed: 0}; + const records = []; + let truncated = false; + const server = http.createServer({ maxHeaderSize: 8192, headersTimeout: 5000, requestTimeout: 10000, + keepAliveTimeout: 1000 }, (request, response) => { + const received = Buffer.from(request.headers.authorization ?? ''); + if (received.length !== authorization.length || !timingSafeEqual(received, authorization)) { + errorReply(response, 401, 'unauthorized'); return; + } + if (request.headers.host !== host || request.headers.origin !== undefined || request.headers.referer !== undefined || + request.method !== 'POST' || request.url !== '/v1/responses' || + !/^application\/json(?:;\s*charset=utf-8)?$/i.test(request.headers['content-type'] ?? '') || + request.headers['content-encoding'] !== undefined || request.headers.expect !== undefined) { + errorReply(response, 400, 'unsupported_request'); return; + } + if (closing || active) { errorReply(response, 503, 'busy'); return; } + if (Number(request.headers['content-length'] ?? 0) > HTTP_BYTES) { errorReply(response, 413, 'request_bound'); return; } + const controller = new AbortController(); + const client = new PrivateConversation(socketPath); + const owner = { controller, client, done: null }; + active = owner; + response.once('close', () => { if (!response.writableFinished) controller.abort(); }); + owner.done = (async () => { + try { + const input = await readBody(request); + if (controller.signal.aborted) throw fail('cancelled'); + const caps = await client.connect(); + check(caps.model_profile === model, 'model_mismatch'); + const conversation = toConversation(input, model, caps); + observed.submitted++; + const started = performance.now(); + const result = await client.submit(conversation, { signal: controller.signal }); + observed.cleanup_confirmed++; + if (result.turn_complete) observed.completed++; else observed.incomplete++; + if (diagnostics) { + if (records.length === 16) truncated = true; + else records.push(Object.freeze({ output_kind: result.output.type, + prompt_tokens: result.prompt_tokens, generated_tokens: result.generated_tokens, + turn_complete: result.turn_complete, + incomplete_reason: result.turn_complete ? null : result.output.reason, + elapsed_ms: Math.floor(performance.now() - started) })); + } + if (controller.signal.aborted || response.destroyed) return; + const events = responseEvents(result); + // No partial model text or tool proposal leaves this process before the + // core's terminal cleanup-confirmed result. SSE is protocol adaptation, + // not a claim of token-by-token backend streaming. + response.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-store', + 'Connection': 'close', 'X-Content-Type-Options': 'nosniff' }); + response.end(events.map(event => `event: ${event.type}\ndata: ${JSON.stringify(event)}\n\n`).join('')); + } catch (error) { + const code = error.message?.startsWith('private_compute_') ? error.code : 'provider_failed'; + errorReply(response, ['busy', 'cleanup_unconfirmed', 'socket_unavailable', 'execution_failed'].includes(code) ? 503 : 400, code); + } finally { + client.close(); + if (active === owner) active = null; + } + })(); + }); + server.maxConnections = 8; + server.maxRequestsPerSocket = 1; + server.on('clientError', (_error, socket) => socket.destroy()); + await new Promise((resolve, reject) => { server.once('error', reject); server.listen(0, '127.0.0.1', resolve); }); + host = `127.0.0.1:${server.address().port}`; + return { baseUrl: `http://${host}/v1`, bearerToken, + // Content-free lifecycle counters for an explicitly started local owner. + get observations() { return Object.freeze({...observed}); }, + // Explicit native trial only. No output text, arguments, commands, IDs or paths. + get diagnostics() { return diagnostics ? Object.freeze({version: 1, + records: Object.freeze([...records]), truncated}) : null; }, + async close() { + closing = true; + const owner = active; + owner?.controller.abort(); + server.closeAllConnections(); + await owner?.done; + await new Promise(resolve => server.close(resolve)); + } }; +} + +module.exports = { startResponsesProvider, toConversation, responseEvents, parseJson }; diff --git a/tests/app-server.test.cjs b/tests/app-server.test.cjs index e1648c4..97b4d06 100644 --- a/tests/app-server.test.cjs +++ b/tests/app-server.test.cjs @@ -78,3 +78,74 @@ test('initialization is single-flight and a close-only stream rejects pending wo const rejected = assert.rejects(pending, /rpc_closed/); destroyed.input.destroy(); await rejected; }); + +test('exact explicit model and writable scope do not broaden default threads', async () => { + const f = fixture({allowedModels: ['qwen3-0.6b-v1'], writableRoot: '/opt/work/project', commandApproval: () => false}); + const init = f.client.initialize(); f.reply(1, {}); await init; + for (const [index, cwd, sandbox] of [[2, '/opt/work/project', 'workspace-write'], [3, '/elsewhere', 'read-only']]) { + const thread = f.client.startThread({model: 'qwen3-0.6b-v1', cwd}); + assert.equal(f.sent.at(-1).params.sandbox, sandbox); + assert.equal(f.sent.at(-1).params.approvalPolicy, 'untrusted'); + f.reply(index, {}); await thread; + } + await assert.rejects(f.client.startThread({model: 'qwen3-8b', cwd: '/opt/work/project'}), /thread_scope/); + f.client.close(); + assert.throws(() => fixture({writableRoot: '/opt/work/project'}), /client_scope/); +}); + +test('optional command policy approves one command only; pre-initialize and other capabilities remain denied', async () => { + let calls = 0; + const f = fixture({commandApproval: params => { calls++; return params.allowed; }}); + const ask = (id, method, params) => f.input.write(JSON.stringify({id, method, params}) + '\n'); + ask('before', 'item/commandExecution/requestApproval', {allowed: true}); + assert.equal(calls, 0); assert.equal(f.sent.at(-1).result.decision, 'decline'); + const init = f.client.initialize(); f.reply(1, {}); await init; + for (const allowed of [true, false, 'acceptForSession', {decision: 'accept'}]) { + ask(String(calls), 'item/commandExecution/requestApproval', {allowed}); + await new Promise(resolve => setImmediate(resolve)); + assert.equal(f.sent.at(-1).result.decision, allowed === true ? 'accept' : 'decline'); + } + ask('file', 'item/fileChange/requestApproval', {allowed: true}); + assert.equal(f.sent.at(-1).result.decision, 'decline'); + ask('token', 'account/chatgptAuthTokens/refresh', {}); + assert.equal(f.sent.at(-1).error.code, -32601); + assert.equal(calls, 4); f.client.close(); +}); + +test('failed or late approval policies cannot grant authority', async () => { + const f = fixture({commandApproval: () => { throw Error('synthetic policy failure'); }}); + const init = f.client.initialize(); f.reply(1, {}); await init; + f.input.write('{"id":"deny","method":"item/commandExecution/requestApproval","params":{}}\n'); + await new Promise(resolve => setImmediate(resolve)); + assert.equal(f.sent.at(-1).result.decision, 'decline'); f.client.close(); + let finish; + const late = fixture({commandApproval: () => new Promise(resolve => { finish = resolve; })}); + const initialized = late.client.initialize(); late.reply(1, {}); await initialized; + late.input.write('{"id":"late","method":"item/commandExecution/requestApproval","params":{}}\n'); + await new Promise(resolve => setImmediate(resolve)); + late.client.close(); finish(true); await new Promise(resolve => setImmediate(resolve)); + assert(!late.sent.some(value => value.id === 'late')); +}); + +test('native execpolicy proposals receive only the one-shot accept or decline decision', async () => { + const {PROJECT, authorize} = require('../src/native-coding-fixture.cjs'); + const f = fixture({writableRoot: PROJECT, commandApproval: params => authorize(params, 'thread', 'turn')}); + try { + const init = f.client.initialize(); f.reply(1, {}); await init; + const params = {kind: 'command', threadId: 'thread', turnId: 'turn', itemId: 'synthetic_read_01', + environmentId: 'local', command: "/bin/bash -c 'python3 -B /opt/fixture.py read'", cwd: PROJECT, + proposedExecpolicyAmendment: ['python3', '-B', '/opt/fixture.py', 'read'], + availableDecisions: ['accept', {acceptWithExecpolicyAmendment: { + execpolicy_amendment: ['python3', '-B', '/opt/fixture.py', 'read']}}, 'cancel']}; + for (const [id, change, decision] of [['observed', {}, 'accept'], + ['other-command', {command: 'python3 -B /opt/other.py read'}, 'decline'], + ['network', {networkApprovalContext: {host: 'private.invalid'}}, 'decline'], + ['other-turn', {turnId: 'old'}, 'decline']]) { + f.input.write(JSON.stringify({id, method: 'item/commandExecution/requestApproval', params: {...params, ...change}})+'\n'); + await new Promise(resolve => setImmediate(resolve)); + assert.deepEqual(f.sent.at(-1), {id, result: {decision}}); + } + assert(!JSON.stringify(f.sent).includes('acceptWithExecpolicyAmendment')); + assert(!JSON.stringify(f.sent).includes('acceptForSession')); + } finally { f.client.close(); } +}); diff --git a/tests/conversation-fixture.cjs b/tests/conversation-fixture.cjs new file mode 100644 index 0000000..9ec3eb5 --- /dev/null +++ b/tests/conversation-fixture.cjs @@ -0,0 +1,67 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Synthetic local protocol fixtures. No model or native Codex process is run here. +'use strict'; +const fs = require('node:fs/promises'); +const net = require('node:net'); +const os = require('node:os'); +const path = require('node:path'); +const { PrivateConversation, expectedLimits, requestLimit } = require('../src/private-conversation.cjs'); + +function caps(model = 'smollm2-360m-v1') { + return { ...expectedLimits(model), execution_slots: 1, max_seconds: 600, + max_request_bytes: requestLimit(model), max_response_bytes: 65536, quarantined: false }; +} +function input(tools = []) { + return { version: 1, visibility: 'private_local', instructions: 'Reply with a truthful bounded proposal.', + history: [{ type: 'message', role: 'user', text: 'Synthetic test input.' }], tools }; +} +function result(output = { type: 'assistant', text: 'Synthetic response; no inference claim.' }, model = 'smollm2-360m-v1') { + return { version: 1, operation: 'compute_private_conversation', model_profile: model, execution_complete: true, + turn_complete: output.type !== 'incomplete', output, prompt_tokens: 10, generated_tokens: 20, + limits: expectedLimits(model), local_only: true, private_data_supported: true, tool_execution: false, + distributed_execution_claimed: false, private_training_claimed: false, model_answer_correctness_proven: false, + cleanup: { complete: true, retained_input: false, retained_report: false } }; +} +function frame(value) { + const body = Buffer.from(JSON.stringify(value)); + const header = Buffer.alloc(4); header.writeUInt32BE(body.length); + return Buffer.concat([header, body]); +} +function reply(socket, request, event, extra = {}) { + socket.write(frame({ version: 1, id: request.id, event, ...extra })); +} +async function fixture(t, handler, capabilities = caps()) { + const directory = await fs.mkdtemp(path.join(os.tmpdir(), 'vp-conversation-')); + await fs.chmod(directory, 0o700); + const socketPath = path.join(directory, 'core.sock'); + const requests = [], sockets = new Set(); + const server = net.createServer(socket => { + sockets.add(socket); + socket.on('error', () => {}); + socket.on('close', () => sockets.delete(socket)); + let data = Buffer.alloc(0); + socket.on('data', chunk => { + data = Buffer.concat([data, chunk]); + while (data.length >= 4 && data.length >= 4 + data.readUInt32BE()) { + const length = data.readUInt32BE(); + const request = JSON.parse(data.subarray(4, 4 + length)); + data = data.subarray(4 + length); + requests.push(request); + if (request.operation.type === 'conversation_capabilities') { + reply(socket, request, 'conversation_capabilities', { capabilities }); + } else handler(socket, request); + } + }); + }); + await new Promise((resolve, reject) => { server.once('error', reject); server.listen(socketPath, resolve); }); + await fs.chmod(socketPath, 0o600); + const client = new PrivateConversation(socketPath); + t.after(async () => { + client.close(); + for (const socket of sockets) socket.destroy(); + await new Promise(resolve => server.close(resolve)); + await fs.rm(directory, { recursive: true }); + }); + return { client, directory, socketPath, requests }; +} +module.exports = { caps, input, result, frame, reply, fixture }; diff --git a/tests/native-coding.test.cjs b/tests/native-coding.test.cjs new file mode 100644 index 0000000..83877b1 --- /dev/null +++ b/tests/native-coding.test.cjs @@ -0,0 +1,335 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Offline policy/helper checks, not native model or coding-loop evidence. +'use strict'; +const {test} = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); +const {spawnSync} = require('node:child_process'); +const {EventEmitter} = require('node:events'); +const {MODEL, PROJECT, PROMPT_SHA256, modelCatalog, runtimeSettings, commandKind, + APPROVAL_DENIALS, approvalDenial, authorize, TASK, CONTINUATION, ITEM_TYPES, + NativeTaskController} = require('../src/native-coding-fixture.cjs'); +const root = path.join(__dirname, '..'); + +function python(source) { + const result = spawnSync('python3', ['-B', '-c', source], {cwd: root, encoding: 'utf8', timeout: 5000}); + assert.equal(result.status, 0, result.stderr); +} + +test('native catalog refuses a replaced prompt; local pinned prompt is retained without truncation', t => { + assert.throws(() => modelCatalog('short replacement instructions'), /native_prompt_mismatch/); + const prompt = path.resolve(root, '..', 'upstream-codex/codex-rs/models-manager/prompt.md'); + if (!fs.existsSync(prompt)) { t.diagnostic('Pinned upstream source absent; positive catalog check not run.'); return; } + const original = fs.readFileSync(prompt, 'utf8'); + const {models: [model]} = modelCatalog(original); + assert.equal(model.slug, MODEL); + assert.equal(model.model_messages.instructions_template, original); + assert.equal(Buffer.byteLength(original), 20903); + assert.equal(model.apply_patch_tool_type, null); + assert.equal(model.supports_reasoning_summary_parameter, false); + assert.equal(model.shell_type, 'unified_exec'); assert.equal(model.tool_mode, 'direct'); + assert.equal(model.input_modalities.length, 1); + assert.equal(model.context_window, 32768); +}); + +test('native configuration has no cloud login, retries, inherited external tools or secret-bearing shell environment', () => { + const settings = runtimeSettings('http://127.0.0.1:12345/v1'); + for (const value of ['model="qwen3-0.6b-v1"', 'model_provider="volparossa"', + 'model_providers.volparossa.requires_openai_auth=false', 'model_providers.volparossa.request_max_retries=0', + 'model_providers.volparossa.stream_max_retries=0', 'features.apps=false', 'features.plugins=false', + 'features.view_image=false', 'model_reasoning_summary="none"', 'mcp_servers={}', + 'shell_environment_policy.exclude=["VOLPAROSSA_PROVIDER_TOKEN"]']) assert(settings.includes(value)); + assert(!settings.some(value => /OPENAI|danger-full-access|approval_policy="never"/.test(value))); + assert.throws(() => runtimeSettings('https://provider.example/v1'), /provider_scope/); + assert.throws(() => runtimeSettings('http://127.0.0.1:12/v1?secret'), /provider_scope/); +}); + +test('exact native shlex commands are recognized without executing or general shell parsing', () => { + const quote = text => "'" + text.replaceAll("'", "'\"'\"'") + "'"; + for (const [text, expected] of [['python3 -B /opt/fixture.py read', 'read'], + ['python3 -B /opt/fixture.py test', 'test'], ["python3 -B /opt/fixture.py edit 'a + b'", 'edit']]) { + assert.equal(commandKind(text), expected); + assert.equal(commandKind(`/bin/bash -c ${quote(text)}`), expected); + assert.equal(commandKind(`/usr/bin/bash -lc ${quote(text)}`), expected); + } + for (const text of ['cat /etc/passwd', 'python3 -B /opt/fixture.py test; id', + 'python3 -B /opt/fixture.py read\nid', '/bin/bash -c "python3 -B /opt/fixture.py test" extra', + "python3 -B /opt/fixture.py edit 'a + b; import os'", "python3 -B /opt/fixture.py edit '$(id)'", + 'python3 -B /opt/other.py read', "python3 -B /opt/fixture.py edit 'a + b", null]) { + assert.equal(commandKind(text), null); + } +}); + +test('fixture approvals bind exact thread, turn, path and one command; no network or policy escalation', () => { + const request = {kind: 'command', threadId: 'thread', turnId: 'turn', itemId: 'item', cwd: PROJECT, + command: '/bin/bash -c "python3 -B /opt/fixture.py read"'}; + assert(authorize(request, 'thread', 'turn')); + for (const [change, reason] of [[{cwd: '/tmp'}, 'cwd'], [{kind: 'writeStdin'}, 'kind'], + [{threadId: 'foreign'}, 'lineage'], [{turnId: 'old'}, 'lineage'], [{itemId: null}, 'item'], + [{itemId: 'x'.repeat(257)}, 'item'], [{command: 'rm -rf /'}, 'command'], + [{networkApprovalContext: {}}, 'network'], [{additionalPermissions: {}}, 'permissions'], + [{proposedNetworkPolicyAmendments: []}, 'network_policy']]) { + assert.equal(authorize({...request, ...change}, 'thread', 'turn'), false); + assert.equal(approvalDenial({...request, ...change}, 'thread', 'turn'), reason); + assert(APPROVAL_DENIALS.includes(reason)); + } + assert.equal(authorize(request, 'thread', null), false); + assert.equal(approvalDenial(null, 'thread', 'turn'), 'lineage'); +}); + +test('observed native read proposal is not a request to persist execution policy', () => { + // Payload shape observed from exact native source 67727e7cf in isolated local + // protocol reproduction, with synthetic Responses and every action declined. + const request = {kind: 'command', threadId: 'thread', turnId: 'turn', itemId: 'synthetic_read_01', + startedAtMs: 1790884427876, environmentId: 'local', + command: "/bin/bash -c 'python3 -B /opt/fixture.py read'", cwd: PROJECT, + commandActions: [{type: 'unknown', command: 'python3 -B /opt/fixture.py read'}], + proposedExecpolicyAmendment: ['python3', '-B', '/opt/fixture.py', 'read'], + availableDecisions: ['accept', {acceptWithExecpolicyAmendment: { + execpolicy_amendment: ['python3', '-B', '/opt/fixture.py', 'read']}}, 'cancel']}; + assert.equal(approvalDenial(request, 'thread', 'turn'), null); + assert.equal(authorize(request, 'thread', 'turn'), true); + // A proposed rule is never authority, even if it would cover another command. + assert.equal(authorize({...request, command: 'python3 -B /opt/other.py read'}, 'thread', 'turn'), false); + assert.equal(authorize({...request, additionalPermissions: {network: true}}, 'thread', 'turn'), false); + assert.equal(authorize({...request, threadId: 'another-owner'}, 'thread', 'turn'), false); +}); + +test('namespace command exposes exact socket/runtime/source inputs, not the host workspace or user home', () => { + python(String.raw` +import runpy +from pathlib import Path +s = runpy.run_path('scripts/smoke_native_coding.py') +c = s['command'](Path('/w/server'), Path('/w/node'), Path('/w/private/socket'), Path('/w/prompt'), + Path('/w/fixture'), '/home/fixture') +for flag in ('--unshare-user','--unshare-net','--unshare-pid','--unshare-ipc','--unshare-uts','--clearenv'): + assert flag in c +assert ['--ro-bind','/','/'] not in [c[i:i+3] for i in range(len(c)-2)] +assert ['--ro-bind','/w/private/socket','/opt/core/compute.sock'] in [c[i:i+3] for i in range(len(c)-2)] +assert not any(c[i:i+2] == ['--setenv',name] for i in range(len(c)-1) + for name in ('HOME','CODEX_HOME','OPENAI_API_KEY')) +assert '--cap-drop' in c and c[c.index('--cap-drop')+1] == 'ALL' +assert c[c.index('--chdir')+1] == '/opt/work/project' +`); +}); + +test('runtime provenance binds binary, source, lock, exact compatibility patch and retained notices', () => { + python(String.raw` +import runpy,tempfile,json,hashlib +from pathlib import Path +s=runpy.run_path('scripts/smoke_native_coding.py') +with tempfile.TemporaryDirectory() as d: + root=Path(d); (root/'third_party').mkdir(); bundle=root/'bundle'; (bundle/'runtime').mkdir(parents=True) + (bundle/'notices').mkdir(); binary=bundle/'runtime/codex-app-server'; binary.write_bytes(b'not executable - offline fixture') + hashes={} + for name in ('LICENSE','NOTICE'): + data=('synthetic '+name).encode(); (bundle/'notices'/name).write_bytes(data) + (bundle/'notices'/name).chmod(0o600); hashes[name]=hashlib.sha256(data).hexdigest() + hashes['codex-rs/Cargo.lock']='a'*64 + pin=dict(revision=s['UPSTREAM'],tree='b'*40,sha256=hashes,patches=[{'file':'synthetic.patch','sha256':'c'*64}]) + (root/'third_party/codex-runtime.json').write_text(json.dumps(pin)) + sha=hashlib.sha256(binary.read_bytes()).hexdigest() + report=dict(version=1,app_server_built=True,staged_source_verified=True,original_source_unchanged=True, + source_revision=pin['revision'],source_tree=pin['tree'],lock_sha256=hashes['codex-rs/Cargo.lock'],local_patches=pin['patches'], + binary=dict(path='runtime/codex-app-server',bytes=binary.stat().st_size,sha256=sha)) + p=bundle/'BUILD_REPORT.json'; p.write_text(json.dumps(report)); p.chmod(0o600) + s['verified_build'].__globals__['ROOT']=root + assert s['verified_build'](str(p),binary,sha)['source_revision']==s['UPSTREAM'] + for change in ({'local_patches':[]},{'app_server_built':False},{'source_revision':'f'*40}): + p.write_text(json.dumps(report|change)) + try:s['verified_build'](str(p),binary,sha) + except ValueError:pass + else:raise AssertionError('unbound runtime accepted') +`); +}); + +test('synthetic helper really reads, applies supplied arithmetic, and runs failing then passing tests', () => { + python(String.raw` +import importlib.util, tempfile, os, sys, contextlib, io, json +from pathlib import Path +spec = importlib.util.spec_from_file_location('fixture','scripts/native_coding_fixture.py') +f = importlib.util.module_from_spec(spec); spec.loader.exec_module(f) +for bad in ('__import__("os")', 'a ** 100000', '[a,b]', 'a; b', 'True', '101', 'secret'): + try: f.expression(bad) + except (ValueError,SyntaxError): pass + else: raise AssertionError('unsafe arithmetic accepted') +with tempfile.TemporaryDirectory() as temporary: + before = Path.cwd() + f.PROJECT = Path(temporary); f.SOURCE = f.PROJECT/'arithmetic.py' + f.SOURCE.write_text(f.ORIGINAL); os.chdir(f.PROJECT) + def action(*args): + sys.argv = ['fixture',*args] + output=io.StringIO() + with contextlib.redirect_stdout(output), contextlib.redirect_stderr(io.StringIO()): status=f.main() + return status,json.loads(output.getvalue()) + try: + assert action('read')[1]['source'] == f.ORIGINAL + assert action('test')[0] == 1 + assert action('edit','a * b')[0] == 0 + assert 'return a * b' in f.SOURCE.read_text() + assert action('test')[0] == 1 + assert action('edit','a + b')[0] == 0 + assert action('test') == (0,{'action':'test','passed':True,'tests':3}) + assert list(f.PROJECT.iterdir()) == [f.SOURCE] + finally: os.chdir(before) +`); +}); + +test('native coding driver requires genuine app-server tools, core terminal cleanup and independent post-edit tests', () => { + const source = fs.readFileSync(path.join(root, 'scripts/smoke_native_coding.cjs'), 'utf8'); + const controller = fs.readFileSync(path.join(root, 'src/native-coding-fixture.cjs'), 'utf8'); + for (const snippet of ['new AppServer(child.stdout, child.stdin', 'await controller.run()', + "spawnSync('/usr/bin/python3', ['-B', '/opt/fixture.py', 'test']", 'report.responses.completed >= 4', + 'report.responses.submitted === report.responses.cleanup_confirmed', 'thread/unsubscribe']) assert(source.includes(snippet)); + assert(!/console\.(log|error)|fake|mockResponse/.test(source)); + assert(source.includes('private_peer_execution_claimed: false, general_coding_quality_claimed: false')); + assert(controller.includes("kind === 'edit' && !report.read")); + assert(controller.includes('item.exitCode === 0')); + assert(source.includes('}, 2400000)')); +}); + +function lifecycle(scripts) { + const report = {read: false, edit: false, test: false, unexpected_command: false, + native_turn_completed: false, accepted_commands: 0, declined_commands: 0, + approval_denials: Object.fromEntries(APPROVAL_DENIALS.map(key => [key, 0])), + turns_started: 0, turns_completed: 0, item_types: Object.fromEntries(ITEM_TYPES.map(key => [key, 0]))}; + const client = new EventEmitter(); + const inputs = [], interrupts = []; + const controller = new NativeTaskController(client, 'thread', report); + client.interrupt = async (...args) => { interrupts.push(args); }; + const request = (turnId, kind) => ({kind: 'command', threadId: 'thread', turnId, itemId: 'synthetic', + cwd: PROJECT, command: kind === 'edit' ? "python3 -B /opt/fixture.py edit 'a + b'" : `python3 -B /opt/fixture.py ${kind}`}); + client.startTurn = async (threadId, text) => { + inputs.push(text); const id = `turn-${inputs.length}`; + client.emit('notification', {method: 'turn/started', params: {threadId, turn: {id}}}); + const item = value => client.emit('notification', {method: 'item/completed', params: {threadId, turnId: id, item: value}}); + const action = kind => { + const proposed = request(id, kind); + const allowed = controller.approve(proposed); + item({...proposed, type: 'commandExecution', status: allowed ? 'completed' : 'declined', exitCode: allowed ? 0 : null}); + return allowed; + }; + const complete = (status = 'completed') => client.emit('notification', { + method: 'turn/completed', params: {threadId, turn: {id, status}}}); + scripts[inputs.length - 1]({action, item, complete, request, id, controller, client}); + return {turn: {id}}; + }; + return {controller, report, inputs, interrupts}; +} + +test('actual controller accepts early started/completed notifications and stops after one complete task', async () => { + const f = lifecycle([({action, item, complete}) => { + action('read'); action('edit'); action('test'); item({type: 'agentMessage', text: 'PRIVATE_CANARY'}); complete(); + }]); + try { await f.controller.run(); } finally { f.controller.dispose(); } + assert.deepEqual(f.inputs, [TASK]); + assert.equal(f.report.turns_started, 1); assert.equal(f.report.turns_completed, 1); + assert.equal(f.report.item_types.commandExecution, 3); assert.equal(f.report.item_types.agentMessage, 1); + assert(!JSON.stringify(f.report).includes('PRIVATE_CANARY')); +}); + +test('one normal incomplete task continues in the same thread without a suggested fix', async () => { + const f = lifecycle([ + ({action, complete}) => { action('read'); complete(); }, + ({action, request, controller, complete}) => { + assert.equal(controller.approve(request('turn-1', 'edit')), false); + assert.equal(controller.report.approval_denials.lineage, 1); + action('edit'); action('test'); complete(); + }, + ]); + try { await f.controller.run(); } finally { f.controller.dispose(); } + assert.deepEqual(f.inputs, [TASK, CONTINUATION]); + assert(!/fixture.py|python|a\s*\+\s*b|exec_command/.test(CONTINUATION)); + assert.equal(f.report.turns_completed, 2); assert.equal(f.report.accepted_commands, 3); +}); + +test('two normally completed but unfinished turns never fabricate success or schedule a third', async () => { + const f = lifecycle([({action, complete}) => { action('read'); complete(); }, + ({complete, controller, request, id}) => { + complete(); + assert.equal(controller.approve(request(id, 'edit')), false, 'ended turn cannot grant late authority'); + }]); + try { await assert.rejects(f.controller.run(), /native_task_incomplete/); } finally { f.controller.dispose(); } + assert.equal(f.inputs.length, 2); assert.equal(f.report.edit, false); assert.equal(f.report.test, false); + assert.equal(f.report.native_turn_completed, true); +}); + +test('failed or interrupted native turns and transport EOF never cause continuation', async () => { + for (const state of ['failed', 'interrupted', 'eof']) { + const f = lifecycle([({complete, client}) => state === 'eof' ? client.emit('closed') : complete(state)]); + try { await assert.rejects(f.controller.run()); } finally { f.controller.dispose(); } + assert.equal(f.inputs.length, 1); assert.equal(f.report.native_turn_completed, false); + } +}); + +test('same approval budget spans both turns and revoked continuation performs no new turn', async () => { + const f = lifecycle([ + ({action, complete}) => { for (let n = 0; n < 5; n++) action('read'); complete(); }, + ({action, complete}) => { assert(action('edit')); assert.equal(action('test'), false); complete(); }, + ]); + try { await assert.rejects(f.controller.run(), /native_task_incomplete/); } finally { f.controller.dispose(); } + assert.equal(f.report.accepted_commands, 6); assert.equal(f.report.approval_denials.budget, 1); + const stopped = lifecycle([({controller}) => { void controller.stop(); }]); + try { await assert.rejects(stopped.controller.run()); } finally { stopped.controller.dispose(); } + assert.equal(stopped.inputs.length, 1); assert.deepEqual(stopped.interrupts, [['thread', 'turn-1']]); +}); + +test('closed receipt rejects invented success and arbitrary text fields', () => { + python(String.raw` +import runpy,tempfile,json,hashlib +from pathlib import Path +s=runpy.run_path('scripts/smoke_native_coding.py') +v=dict(version=1,kind='native-codex-core-coding',success=False,phase='native-turn',model='qwen3-0.6b-v1', + full_native_prompt_sha256=s['PROMPT_SHA256'],before_sha256=hashlib.sha256(s['ORIGINAL']).hexdigest(),after_sha256=None, + native_turn_completed=False,read=False,edit=False,test=False,independent_test_passed=False, + accepted_commands=0,declined_commands=0,unexpected_command=False,thread_unsubscribed=False,responses=None, + private_peer_execution_claimed=False,general_coding_quality_claimed=False,runtime_exit=1,forced_stop=False, + diagnostic='native_coding_incomplete') +with tempfile.TemporaryDirectory() as d: + p=Path(d)/'receipt.json'; p.write_text(json.dumps(v)); assert s['closed_receipt'](p)==v + for changes in ({'success':True},{'diagnostic':'raw private command'},{'full_native_prompt_sha256':'wrong'}, + {'url':'https://private.invalid'},{'accepted_commands':999}): + p.write_text(json.dumps(v|changes)) + try:s['closed_receipt'](p) + except ValueError:pass + else:raise AssertionError('unproven receipt accepted') + denials={reason:0 for reason in s['APPROVAL_DENIALS']}; denials['command']=1 + v2=v|dict(version=2,declined_commands=1,approval_denials=denials) + p.write_text(json.dumps(v2)); assert s['closed_receipt'](p)==v2 + for changes in ({'version':3},{'version':True},{'version':1},{'declined_commands':0}, + {'approval_denials':{}},{'approval_denials':denials|{'raw private command':1}}, + {'approval_denials':denials|{'command':True}}, + {'approval_denials':denials|{'command':-1}}, + {'approval_denials':denials|{'command':17}}, + {'approval_denials':denials|{'command':'private command'}}): + p.write_text(json.dumps(v2|changes)) + try:s['closed_receipt'](p) + except ValueError:pass + else:raise AssertionError('invalid denial counters accepted') + p.write_text(json.dumps({k:value for k,value in v2.items() if k!='approval_denials'})) + try:s['closed_receipt'](p) + except ValueError:pass + else:raise AssertionError('v2 counters missing') + v3=v2|dict(version=3,turns_started=2,turns_completed=2, + item_types=dict(commandExecution=1,agentMessage=1,userMessage=2,reasoning=0,other=0), + responses=dict(submitted=1,completed=1,incomplete=0,cleanup_confirmed=1), + response_diagnostics=dict(version=1,truncated=False,records=[dict(output_kind='assistant', + prompt_tokens=9000,generated_tokens=100,turn_complete=True,incomplete_reason=None,elapsed_ms=300000)])) + p.write_text(json.dumps(v3)); assert s['closed_receipt'](p)==v3 + for change in ({'turns_started':3},{'turns_completed':True},{'response_diagnostics':None}, + {'item_types':v3['item_types']|{'private':'PRIVATE_CANARY'}}): + p.write_text(json.dumps(v3|change)) + try:s['closed_receipt'](p) + except ValueError:pass + else:raise AssertionError('unbounded diagnostics accepted') + for change in ({'output_kind':'PRIVATE_CANARY'},{'text':'PRIVATE_CANARY'},{'prompt_tokens':12289}, + {'generated_tokens':1025},{'elapsed_ms':3600001},{'turn_complete':1}): + bad=v3|dict(response_diagnostics=v3['response_diagnostics']|dict(records=[v3['response_diagnostics']['records'][0]|change])) + p.write_text(json.dumps(bad)) + try:s['closed_receipt'](p) + except ValueError:pass + else:raise AssertionError('private or unbounded diagnostic accepted') +`); + assert.match(PROMPT_SHA256, /^[a-f0-9]{64}$/); +}); diff --git a/tests/private-conversation.test.cjs b/tests/private-conversation.test.cjs new file mode 100644 index 0000000..5d24493 --- /dev/null +++ b/tests/private-conversation.test.cjs @@ -0,0 +1,95 @@ +// SPDX-License-Identifier: GPL-3.0-only +'use strict'; +const assert = require('node:assert/strict'); +const fs = require('node:fs/promises'); +const { test } = require('node:test'); +const { validateConversation, PrivateConversation } = require('../src/private-conversation.cjs'); +const { caps, input, result, frame, reply, fixture } = require('./conversation-fixture.cjs'); + +test('actual framed socket uses separate handshake and yields exact cleanup-confirmed conversation result', async t => { + const original = result(); + const f = await fixture(t, (socket, request) => { + const bytes = Buffer.concat([frame({ version: 1, id: request.id, event: 'admitted' }), + frame({ version: 1, id: request.id, event: 'result', result: original })]); + socket.write(bytes.subarray(0, 2)); + setImmediate(() => socket.write(bytes.subarray(2))); + }); + assert.deepEqual(await f.client.connect(), caps()); + assert.deepEqual(await f.client.submit(input()), original); + assert.deepEqual(f.requests.map(row => row.operation.type), ['conversation_capabilities', 'submit_conversation']); + await assert.rejects(f.client.ask({ question: 'legacy', context: 'must not send' })); +}); + +test('public/cloud/widened/unknown capabilities fail before any task', async t => { + for (const change of [{ network_access: true }, { training: true }, { cloud_fallback: true }, + { max_prompt_tokens: 999999 }, { native_tool_template: true }, { model_profile: 'unknown' }, + { arbitrary_json_schema_validation: true }, { max_seconds: 0 }, { max_request_bytes: 999999 }]) { + const f = await fixture(t, () => assert.fail('unexpected submit'), { ...caps(), ...change }); + await assert.rejects(f.client.connect()); + assert.equal(f.requests.length, 1); + } +}); + +test('history preserves exact namespaced calls and requires one matching result before the next message', () => { + const value = input([{ type: 'function', namespace: 'workspace', name: 'read', description: 'Read approved file.', + parameters: { type: 'object', properties: { file: { type: 'string' } } } }]); + value.history.push({ type: 'function_call', namespace: 'workspace', name: 'read', call_id: 'c1', arguments: { file: 'demo' } }, + { type: 'tool_result', call_id: 'c1', output: 'Public fixture bytes.' }); + validateConversation(value); + for (const alter of [v => { v.history[2].call_id = 'other'; }, v => { v.history[1].namespace = null; }, + v => { v.history.splice(2, 0, { type: 'message', role: 'user', text: 'interrupt' }); }, + v => { v.history.push(v.history[2]); }, v => { v.tools = []; }, + v => { v.history[1].arguments = 'not object'; }, v => { v.instructions = 'x'.repeat(4097); }]) { + const invalid = structuredClone(value); alter(invalid); + assert.throws(() => validateConversation(invalid)); + } +}); + +test('cleanup uncertainty, wrong tool identity and historical call reuse never expose output', async t => { + const request = input([{ type: 'custom', name: 'patch', namespace: null, description: 'Propose a literal patch.' }]); + request.history.push({ type: 'custom_tool_call', name: 'patch', namespace: null, call_id: 'old', input: 'example' }, + { type: 'tool_result', call_id: 'old', output: 'fixture' }); + for (const output of [ + { type: 'custom_tool_call', name: 'other', namespace: null, call_id: 'new', input: 'must not run' }, + { type: 'custom_tool_call', name: 'patch', namespace: null, call_id: 'old', input: 'must not run' }, + { type: 'custom_tool_call', name: 'patch', call_id: 'new', input: 'missing namespace' }, + ]) { + const f = await fixture(t, (socket, message) => { + reply(socket, message, 'admitted'); reply(socket, message, 'result', { result: result(output) }); + }); + await f.client.connect(); await assert.rejects(f.client.submit(request)); + } + const invalid = result(); invalid.cleanup.complete = false; + const f = await fixture(t, (socket, message) => { + reply(socket, message, 'admitted'); reply(socket, message, 'result', { result: invalid }); + }); + await f.client.connect(); await assert.rejects(f.client.submit(input()), { code: 'cleanup_unconfirmed' }); +}); + +test('cancellation acknowledgement does not settle until the exact task cleanup terminal', async t => { + let task, authorize; + const ready = new Promise(resolve => { authorize = resolve; }); + const f = await fixture(t, (socket, request) => { + if (request.operation.type === 'submit_conversation') { task = request; reply(socket, request, 'admitted'); } + else { assert.equal(request.operation.task_id, task.id); + reply(socket, request, 'cancel_requested', { task_id: task.id }); + authorize(() => reply(socket, task, 'error', { code: 'cancelled' })); } + }); + await f.client.connect(); + const controller = new AbortController(); + const pending = f.client.submit(input(), { signal: controller.signal }); + const rejected = assert.rejects(pending, { code: 'cancelled' }); + controller.abort(); + const finish = await ready; + let settled = false; pending.catch(() => { settled = true; }); + await new Promise(resolve => setImmediate(resolve)); assert.equal(settled, false); + finish(); await rejected; +}); + +test('conversation inherits exact owned socket/parent boundary and rejects false result correlation', async t => { + const f = await fixture(t, (socket, request) => reply(socket, { id: 'f'.repeat(32) }, 'result', { result: result() })); + await fs.chmod(f.socketPath, 0o666); + await assert.rejects(new PrivateConversation(f.socketPath).connect(), { code: 'socket_ownership' }); + await fs.chmod(f.socketPath, 0o600); + await f.client.connect(); await assert.rejects(f.client.submit(input()), { code: 'invalid_response' }); +}); diff --git a/tests/responses-provider.test.cjs b/tests/responses-provider.test.cjs new file mode 100644 index 0000000..d0564f8 --- /dev/null +++ b/tests/responses-provider.test.cjs @@ -0,0 +1,266 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Real HTTP + Unix sockets, synthetic protocol peer; not model/native coding proof. +'use strict'; +const assert = require('node:assert/strict'); +const http = require('node:http'); +const { test } = require('node:test'); +const { startResponsesProvider, toConversation, parseJson } = require('../src/responses-provider.cjs'); +const { caps, input, result, reply, fixture } = require('./conversation-fixture.cjs'); + +function request(model = 'smollm2-360m-v1') { + return { model, instructions: 'Answer with a bounded truthful proposal.', + input: [{ type: 'message', role: 'user', content: [{ type: 'input_text', text: 'Synthetic fixture.' }] }], + tools: [], tool_choice: 'auto', parallel_tool_calls: true, reasoning: {}, store: false, stream: true, + include: ['reasoning.encrypted_content'], prompt_cache_key: 'fixture-no-cache', text: null, + client_metadata: { session_id: 'fixture-no-telemetry', thread_id: 'fixture' } }; +} +async function start(t, handler, model = 'smollm2-360m-v1', diagnostics = false) { + const f = await fixture(t, handler, caps(model)); + const provider = await startResponsesProvider({ socketPath: f.socketPath, model, diagnostics }); + // Register explicitly in the parent fixture cleanup order. Closing the HTTP + // listener is awaited in each test before the synthetic Unix backend closes. + return { ...f, provider }; +} +function send(provider, value, headers = {}) { + const body = typeof value === 'string' ? value : JSON.stringify(value); + return new Promise((resolve, reject) => { + const req = http.request(`${provider.baseUrl}/responses`, { method: 'POST', + headers: { 'content-type': 'application/json', authorization: `Bearer ${provider.bearerToken}`, + 'content-length': Buffer.byteLength(body), ...headers } }, response => { + let data = ''; + response.on('data', chunk => { data += chunk; }); + response.on('end', () => resolve({ status: response.statusCode, headers: response.headers, body: data })); + }); + req.on('error', reject); req.end(body); + }); +} +function events(response) { + assert.equal(response.status, 200); + assert.equal(response.headers['content-type'], 'text/event-stream'); + const rows = response.body.trim().split('\n\n').map(block => { + const [type, data] = block.split('\n'); + const value = JSON.parse(data.slice(6)); + assert.equal(type, `event: ${value.type}`); + return value; + }); + assert.deepEqual(rows.map(value => value.sequence_number), rows.map((_value, index) => index)); + return rows; +} + +test('explicit loopback provider authenticates before IPC and retains no default/cloud listener', async t => { + const f = await start(t, () => assert.fail('no generation expected')); + try { + assert.match(f.provider.baseUrl, /^http:\/\/127\.0\.0\.1:\d+\/v1$/); + assert.equal(f.provider.bearerToken.length, 43); + assert.equal(f.requests.length, 0); + assert.equal(f.provider.diagnostics, null); + for (const headers of [{ authorization: 'Bearer wrong' }, { origin: 'https://untrusted.invalid' }, + { referer: 'https://untrusted.invalid/' }, { host: 'attacker.invalid' }, { 'content-encoding': 'gzip' }]) { + const response = await send(f.provider, request(), headers); + assert.ok([400, 401].includes(response.status)); + } + assert.equal(f.requests.length, 0); + } finally { await f.provider.close(); } +}); + +test('native-shaped request streams only the exact final answer after cleanup, with real token counts', async t => { + let finish, admitted; + const ready = new Promise(resolve => { admitted = resolve; }); + const original = result(); + const f = await start(t, (socket, message) => { + reply(socket, message, 'admitted'); + finish = () => reply(socket, message, 'result', { result: original }); + admitted(); + }); + try { + let settled = false; + const pending = send(f.provider, request()); pending.then(() => { settled = true; }); + await ready; + await new Promise(resolve => setImmediate(resolve)); assert.equal(settled, false); + assert.deepEqual(f.provider.observations, {submitted: 1, completed: 0, incomplete: 0, cleanup_confirmed: 0}); + assert.deepEqual(f.requests.map(row => row.operation.type), ['conversation_capabilities', 'submit_conversation']); + const forwarded = f.requests[1].operation.conversation; + assert.equal(forwarded.instructions, request().instructions); + assert.equal(forwarded.history[0].text, 'Synthetic fixture.'); + assert.ok(!JSON.stringify(forwarded).includes('fixture-no-cache')); + assert.ok(!JSON.stringify(forwarded).includes('fixture-no-telemetry')); + finish(); + const response = await pending; + const rows = events(response); + assert.equal(rows[0].type, 'response.created'); + assert.equal(rows.at(-1).type, 'response.completed'); + assert.equal(rows.at(-1).response.output[0].content[0].text, original.output.text); + assert.deepEqual(rows.at(-1).response.usage, { input_tokens: 10, output_tokens: 20, total_tokens: 30, + input_tokens_details: { cached_tokens: 0 }, output_tokens_details: { reasoning_tokens: 0 } }); + assert.equal(rows.at(-1).response.end_turn, true); + assert.deepEqual(f.provider.observations, {submitted: 1, completed: 1, incomplete: 0, cleanup_confirmed: 1}); + assert(Object.isFrozen(f.provider.observations)); + const followup = request(); + followup.input.push(rows.at(-1).response.output[0], { type: 'message', role: 'user', content: 'Continue.' }); + assert.equal(toConversation(followup, followup.model, caps()).history[1].text, original.output.text); + } finally { await f.provider.close(); } +}); + +test('namespaced function and literal custom proposals preserve IDs and never execute tools', async t => { + for (const custom of [false, true]) { + const proposed = { type: custom ? 'custom_tool_call' : 'function_call', call_id: 'fresh-call', + name: custom ? 'patch' : 'read', namespace: 'workspace', + ...(custom ? { input: 'literal synthetic patch; do not execute' } : { arguments: { file: 'demo.rs' } }) }; + const f = await start(t, (socket, message) => { + reply(socket, message, 'admitted'); reply(socket, message, 'result', { result: result(proposed) }); + }); + try { + const body = request(); + body.tools = [{ type: 'namespace', name: 'workspace', description: 'Only owner-approved inputs.', tools: [ + { type: custom ? 'custom' : 'function', name: proposed.name, description: 'Fixture tool.', + ...(custom ? { format: { type: 'text' } } : { strict: false, parameters: { type: 'object' } }) }, + ] }]; + const rows = events(await send(f.provider, body)); + const answer = rows.at(-1).response.output[0]; + assert.equal(answer.type, proposed.type); assert.equal(answer.namespace, proposed.namespace); + assert.equal(answer.call_id, proposed.call_id); + assert.equal(rows.at(-1).response.end_turn, false); + if (custom) assert.equal(answer.input, proposed.input); + else assert.deepEqual(JSON.parse(answer.arguments), proposed.arguments); + const history = structuredClone(body); + history.input.push(answer, { type: custom ? 'custom_tool_call_output' : 'function_call_output', + call_id: answer.call_id, output: [{ type: 'input_text', text: 'user-approved fixture result' }] }); + const mapped = toConversation(history, body.model, caps()); + assert.equal(mapped.history.at(-1).call_id, answer.call_id); + assert.match(mapped.tools[0].description, /Only owner-approved inputs/); + assert.equal(f.requests.length, 2); + } finally { await f.provider.close(); } + } +}); + +test('all incomplete core results emit response.incomplete without output or completed event', async t => { + for (const reason of ['token_limit', 'wire_truncated', 'invalid_output']) { + const f = await start(t, (socket, message) => { + reply(socket, message, 'admitted'); reply(socket, message, 'result', { result: result({ type: 'incomplete', reason }) }); + }, 'smollm2-360m-v1', true); + try { + const rows = events(await send(f.provider, request())); + assert.deepEqual(rows.map(row => row.type), ['response.created', 'response.incomplete']); + assert.equal(rows.at(-1).response.status, 'incomplete'); + assert.deepEqual(rows.at(-1).response.output, []); + assert.notEqual(rows.at(-1).response.incomplete_details.reason, 'interrupted'); + const diagnostic = f.provider.diagnostics.records[0]; + assert.equal(diagnostic.output_kind, 'incomplete'); assert.equal(diagnostic.incomplete_reason, reason); + assert.equal(diagnostic.turn_complete, false); + } finally { await f.provider.close(); } + } +}); + +test('opt-in diagnostics are content-free immutable and bounded across real HTTP replies', async t => { + const f = await start(t, (socket, message) => { + reply(socket, message, 'admitted'); + reply(socket, message, 'result', {result: result({type: 'assistant', text: 'PRIVATE_CANARY'})}); + }, 'smollm2-360m-v1', true); + try { + for (let n = 0; n < 17; n++) events(await send(f.provider, request())); + const value = f.provider.diagnostics; + assert.equal(value.records.length, 16); assert.equal(value.truncated, true); + assert(Object.isFrozen(value)); assert(Object.isFrozen(value.records)); + for (const row of value.records) { + assert(Object.isFrozen(row)); + assert.deepEqual(Object.keys(row).sort(), ['elapsed_ms', 'generated_tokens', 'incomplete_reason', + 'output_kind', 'prompt_tokens', 'turn_complete']); + assert.equal(row.output_kind, 'assistant'); assert.equal(row.incomplete_reason, null); + assert.equal(row.turn_complete, true); assert.equal(row.prompt_tokens, 10); assert.equal(row.generated_tokens, 20); + assert(Number.isInteger(row.elapsed_ms) && row.elapsed_ms >= 0 && row.elapsed_ms <= 5000); + } + assert(!JSON.stringify(value).includes('PRIVATE_CANARY')); + assert.equal(f.provider.observations.completed, 17); + } finally { await f.provider.close(); } +}); + +test('invalid history, grammar, reasoning and unknown features reject without a submit or truncation', async t => { + const f = await start(t, () => assert.fail('must not submit')); + try { + for (const change of [ + { previous_response_id: 'hidden-state' }, { reasoning: { effort: 'high' } }, { store: true }, + { input: [{ type: 'input_image', image_url: 'private' }] }, { instructions: 'x'.repeat(4097) }, + { tools: [{ type: 'function', name: 'bad', strict: true, description: 'No claimed validation.', parameters: { type: 'object' } }] }, + { tools: [{ type: 'custom', name: 'bad', description: 'No grammar.', format: { type: 'grammar', syntax: 'lark', definition: 'bad' } }] }, + { input: [{ type: 'function_call_output', call_id: 'missing', output: 'private' }] }, + { input: [{ type: 'message', role: 'user', content: [{ type: 'input_text', text: 'x' }, { type: 'input_image', image_url: 'private' }] }] }, + ]) { + const response = await send(f.provider, { ...request(), ...change }); + assert.equal(response.status, 400); + assert.ok(!response.body.includes('private')); + } + assert.ok(f.requests.every(row => row.operation.type === 'conversation_capabilities')); + } finally { await f.provider.close(); } +}); + +test('duplicate keys, unsafe numbers and depth overflow are rejected explicitly', () => { + for (const value of ['{"same":1,"same":2}', '{"same":1,"\\u0073ame":2}', + '{"value":9007199254740993}', '{"value":1e999}', '['.repeat(34) + '0' + ']'.repeat(34)]) { + assert.throws(() => parseJson(value)); + } + assert.deepEqual(parseJson('{"quoted":"{not: syntax}","nested":[null,true,1.25]}'), + { quoted: '{not: syntax}', nested: [null, true, 1.25] }); +}); + +test('unknown cleanup prevents any SSE and returns a closed error without raw model text', async t => { + const invalid = result({ type: 'assistant', text: 'private-canary-never-export' }); + invalid.cleanup.complete = false; + const f = await start(t, (socket, message) => { + reply(socket, message, 'admitted'); reply(socket, message, 'result', { result: invalid }); + }); + try { + const response = await send(f.provider, request()); + assert.equal(response.status, 503); + assert.equal(JSON.parse(response.body).error.code, 'cleanup_unconfirmed'); + assert.ok(!response.body.includes('private-canary')); + assert.ok(!response.body.includes('response.completed')); + } finally { await f.provider.close(); } +}); + +test('HTTP disconnect cancels exact core task, holds single slot until terminal cleanup', async t => { + let admitted, acknowledge, task; + const started = new Promise(resolve => { admitted = resolve; }); + const cancelled = new Promise(resolve => { acknowledge = resolve; }); + const f = await start(t, (socket, message) => { + if (message.operation.type === 'submit_conversation') { task = message; reply(socket, message, 'admitted'); admitted(); } + else { + assert.equal(message.operation.task_id, task.id); + reply(socket, message, 'cancel_requested', { task_id: task.id }); + acknowledge(() => reply(socket, task, 'error', { code: 'cancelled' })); + } + }); + try { + const body = JSON.stringify(request()); + const req = http.request(`${f.provider.baseUrl}/responses`, { method: 'POST', headers: { + authorization: `Bearer ${f.provider.bearerToken}`, 'content-type': 'application/json', 'content-length': Buffer.byteLength(body), + } }); + req.on('error', () => {}); req.end(body); + await started; req.destroy(); + const finish = await cancelled; + assert.equal((await send(f.provider, request())).status, 503); + finish(); + await new Promise(resolve => setTimeout(resolve, 10)); + assert.deepEqual(f.requests.map(row => row.operation.type), ['conversation_capabilities', 'submit_conversation', 'cancel']); + } finally { await f.provider.close(); } +}); + +test('larger Qwen profile transmits full instructions beyond legacy frame without widening Q&A', async t => { + const model = 'qwen3-0.6b-v1'; + const f = await start(t, (socket, message) => { + reply(socket, message, 'admitted'); reply(socket, message, 'result', { result: result(undefined, model) }); + }, model); + try { + const body = request(model); body.instructions = 'Bounded native fixture instruction. '.repeat(1500); + body.input.unshift({ type: 'message', role: 'developer', content: [ + { type: 'input_text', text: 'Exact developer instructions.' }, + { type: 'input_text', text: 'Keep private workspace context private.' }, + ] }); + assert.ok(Buffer.byteLength(JSON.stringify(body)) > 32768); + const rows = events(await send(f.provider, body)); + assert.equal(rows.at(-1).response.model, model); + assert.equal(f.requests[1].operation.conversation.instructions, body.instructions); + assert.deepEqual(f.requests[1].operation.conversation.history[0], { type: 'message', role: 'developer', + text: 'Exact developer instructions.\n\nKeep private workspace context private.' }); + assert.throws(() => toConversation({ ...body, model: 'smollm2-360m-v1' }, 'smollm2-360m-v1', caps())); + } finally { await f.provider.close(); } +}); diff --git a/tests/test_build_codex_runtime.py b/tests/test_build_codex_runtime.py index ef648f1..2238efc 100644 --- a/tests/test_build_codex_runtime.py +++ b/tests/test_build_codex_runtime.py @@ -4,8 +4,10 @@ import importlib.util import json from pathlib import Path +import subprocess import tempfile import unittest +from unittest.mock import patch SCRIPT = Path(__file__).resolve().parents[1] / 'scripts/build_codex_runtime.py' SPEC = importlib.util.spec_from_file_location('builder', SCRIPT) @@ -14,6 +16,89 @@ class BuildBoundaries(unittest.TestCase): + def test_only_fetch_resolves_and_binds_one_resolver_file(self): + state = Path('/private/build') + with patch.object(BUILDER.Path, 'resolve', side_effect=AssertionError('offline resolved DNS')): + offline = BUILDER.sandbox_command(['cargo', 'build', '--offline'], state) + self.assertIn('--unshare-net', offline) + with tempfile.TemporaryDirectory() as name: + resolver = Path(name) / 'resolv.conf' + resolver.write_text('nameserver 192.0.2.53\n') + with patch.object(BUILDER.Path, 'resolve', return_value=resolver) as resolve: + online = BUILDER.sandbox_command(['cargo', 'fetch', '--locked'], state, network=True) + resolve.assert_called_once_with(strict=True) + self.assertEqual(online[-7:], ['--ro-bind', str(resolver), str(resolver), '--', + 'cargo', 'fetch', '--locked']) + self.assertNotIn('--unshare-net', online) + self.assertIn('--tmpfs', online) + resolver.write_bytes(b'x' * (64 * 1024 + 1)) + with patch.object(BUILDER.Path, 'resolve', return_value=resolver): + with self.assertRaises(ValueError): + BUILDER.sandbox_command(['cargo', 'fetch'], state, network=True) + + @unittest.skipUnless(Path('/usr/bin/bwrap').is_file(), 'requires disposable bubblewrap namespaces') + def test_real_synthetic_run_resolver_is_read_only_and_fetch_only(self): + # Neither sandbox has Internet access. The outer namespace supplies a + # fake /etc -> /run resolver; the inner uses the actual builder helper. + with tempfile.TemporaryDirectory() as name: + root = Path(name) + (root / 'etc').mkdir() + (root / 'etc/resolv.conf').symlink_to('/run/build-resolver-test/resolv.conf') + (root / 'resolver').write_text('nameserver 192.0.2.53\n') + (root / 'private').write_text('unrelated-runtime-state') + state = root / 'build' + (state / 'source/codex-rs').mkdir(parents=True) + probe = ''' +import errno, os, sys +from pathlib import Path +resolver = Path('/etc/resolv.conf') +assert resolver.is_symlink() +assert not list(Path('/home').iterdir()) and not list(Path('/root').iterdir()) +assert not Path('/run/build-private-test/token').exists() +if sys.argv[1] == 'online': + assert resolver.read_text() == 'nameserver 192.0.2.53\\n' + assert set(str(p) for p in Path('/run').rglob('*')) == { + '/run/build-resolver-test', '/run/build-resolver-test/resolv.conf'} + try: + resolver.open('w') + except OSError as error: + assert error.errno == errno.EROFS + else: + raise AssertionError('resolver is writable') +else: + assert not resolver.exists() and not list(Path('/run').iterdir()) + assert os.readlink('/proc/self/ns/net') != sys.argv[2] +try: + Path('forbidden-write').write_text('x') +except OSError as error: + assert error.errno == errno.EROFS +else: + raise AssertionError('source is writable') +''' + driver = ''' +import importlib.util, os, subprocess, sys +from pathlib import Path +spec = importlib.util.spec_from_file_location('builder', sys.argv[1]) +builder = importlib.util.module_from_spec(spec) +spec.loader.exec_module(builder) +for network in (True, False): + command = builder.sandbox_command(['/usr/bin/python3', '-I', '-c', sys.argv[3], + 'online' if network else 'offline', os.readlink('/proc/self/ns/net')], + Path(sys.argv[2]), network=network) + subprocess.run(command, check=True, timeout=15, env={'PATH': '/usr/bin:/bin', 'LANG': 'C.UTF-8'}) +''' + command = ['/usr/bin/bwrap', '--die-with-parent', '--unshare-net', '--ro-bind', '/', '/', + '--tmpfs', '/run', '--ro-bind', str(root / 'etc'), '/etc', + '--ro-bind', str(root / 'resolver'), '/run/build-resolver-test/resolv.conf', + '--ro-bind', str(root / 'private'), '/run/build-private-test/token', + '--proc', '/proc', '--dev', '/dev', '--', '/usr/bin/python3', '-I', '-c', + driver, str(SCRIPT), str(state), probe] + subprocess.run(command, check=True, timeout=40, + env={'PATH': '/usr/bin:/bin', 'LANG': 'C.UTF-8'}) + self.assertEqual((root / 'resolver').read_text(), 'nameserver 192.0.2.53\n') + self.assertEqual((root / 'private').read_text(), 'unrelated-runtime-state') + self.assertFalse((state / 'source/codex-rs/forbidden-write').exists()) + def test_paths(self): for value in ('../escape', '/absolute', 'a/../b', 'a//b', 'x\\y', 'x\ny'): with self.assertRaises(ValueError):