Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 15 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,25 +26,32 @@ automatically public training or cache material.

## First executable slice

The extension implements two explicit commands:
The extension implements explicit commands:

- **VOLPAROSSA: Ask About Selected Code (Private, Local)** sends only a confirmed
question and selection to an existing same-owner `compute private-serve` socket.
Responses appear as untrusted plaintext; no changes are applied automatically.
- **VOLPAROSSA: Show Compute Capabilities** queries that service without sending
code or claiming that a model has successfully executed.
- **VOLPAROSSA: Run Native Coding Task (Private, Local)** explicitly launches a
prepared, source-verified open Codex runtime in an isolated Linux workspace and
connects it to the existing local VOLPAROSSA conversation service. The native
agent can read, change and check that selected project, with one-shot command
approvals and cancellation. See [setup and current proof limits](docs/NATIVE_EDITOR.md).

The current core interface permits **512 UTF-8 bytes for the question and 4096
The selected-code advice interface permits **512 UTF-8 bytes for the question and 4096
for the selection**, subject to the selected model's smaller token budget.
Over-limit inputs fail instead of being silently shortened. Partial model output
remains labeled partial. Cancellation is forwarded; uncertain cleanup is not
reported as success. There is no public-peer or OpenAI fallback.

These first commands use the core directly. They are **not yet routed through
Codex**. The separate app-server client implements the pinned NDJSON handshake,
The first two commands use the core directly, not through Codex. The new native
coding command uses the app-server, but is **not yet proved in a native editor
with real model-driven editing**. The app-server client implements the pinned NDJSON handshake,
thread/turn requests, notifications and interruption, and declines tool approvals
by default. An explicit caller can supply a narrowly scoped per-command approval
policy; the normal extension does not enable it. Its focused protocol tests are
policy; only the explicit native coding command enables an interactive one-shot
policy for the selected project. Its focused protocol tests are
now complemented by a **real, source-built app-server lifecycle trial**:
initialization, an ephemeral VOLPAROSSA-provider thread, exact unsubscribe and
clean shutdown pass in disposable namespaces without OpenAI credentials or
Expand Down Expand Up @@ -97,9 +104,9 @@ npm run check

- Prove the new conversation/provider interface with an actual model and the
native Codex Responses/tool loop; the bounded Q&A endpoint stays separate.
- Connect the built runtime to the extension and core provider, retaining its
isolated configuration and upstream notices; never use the owner's OpenAI login
or cloud fallback.
- Prove the new explicit runtime/extension/provider connection in a native editor,
retaining its isolated configuration and upstream notices; never use the
owner's OpenAI login or cloud fallback.
- Complete native conversation/tool interoperability, reviewable diffs and local
approvals, then prove an actual edit-and-test coding task end to end.
- Delegate eligible work through the core's cooperative scheduler, with explicit
Expand Down
17 changes: 17 additions & 0 deletions docs/NATIVE_CODING_TRIAL.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,6 +157,23 @@ arguments, paths or identifiers. Historical version-1/2 receipts remain readable
The bounded continuation and these diagnostics have offline controller/protocol
coverage, not a newly successful model-driven read/edit/test proof.

The actual [run 36932657647](https://github.com/VOLPAROSSA/volparossa/actions/runs/36932657647)
on core `c3fb587f6cdcdc9fd1e0a1dd31a9a0bb6001706c` / Code `2f7014b0`
now reaches that continuation: one read is executed, the first native turn ends,
then another real tool proposal is refused by the fixture's exact command policy.
One command is accepted, one declined in category `command`; edit and test remain
false. All three model responses are complete and cleanup-confirmed (function
call, assistant, function call). The rejected command text is not retained, so
its intended action and correctness are unknown. Runtime exits normally and
private/service cleanup and unchanged host-state checks pass. This is a failed
read/edit/test proof, not a successful coding task or a reason to loosen that
fixture's existing success criteria.

The separate [native editor integration](NATIVE_EDITOR.md) uses the user's actual
chosen workspace and interactive command approvals, rather than the arithmetic
fixture's special command list. Its frontend/launcher implementation and narrow
checks do not supersede this failed model-driven evidence.

Success requires the actual app-server's command-completion events, changed file
hash, independent passing tests, at least four cleanup-confirmed real core
responses, exact thread unsubscribe and graceful runtime exit. Partial responses,
Expand Down
196 changes: 196 additions & 0 deletions docs/NATIVE_EDITOR.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,196 @@
# Native coding in the editor

The explicit **VOLPAROSSA: Run Native Coding Task (Private, Local)** command
connects the editor to the source-built open Codex app-server and the existing
VOLPAROSSA conversation service. The runtime may read and edit the chosen project
and execute approved tools. The extension does not contain a model or a second
peer scheduler, and does not supply a predefined repair or fabricate tool output.

This is a Linux development integration. Controller and launcher checks are not
proof of reliable model-driven coding or a native VS Code/VSCodium end-to-end
trial. The original selected-code advice commands remain available separately.

## Explicit setup

Prepare the [pinned runtime](RUNTIME_BUILD.md) and an existing same-owner private
VOLPAROSSA conversation service using the `qwen3-0.6b-v1` profile. Nothing is
downloaded or installed by this command. In **user settings**, configure:

```json
{
"volparossaCode.privateSocket": "/absolute/private-directory/compute.sock",
"volparossaCode.nativeRuntime": {
"version": 1,
"appServer": "/absolute/runtime-bundle/runtime/codex-app-server",
"appServerSha256": "<SHA-256 from the verified build report>",
"buildReport": "/absolute/runtime-bundle/BUILD_REPORT.json",
"node": "/absolute/prepared-node/bin/node",
"nodeSha256": "<SHA-256 of the explicitly prepared Node executable>",
"upstreamPrompt": "/absolute/pinned-source/codex-rs/models-manager/prompt.md"
}
}
```

Use canonical absolute paths and lowercase 64-character SHA-256 values, not the
placeholders above. The launcher verifies the exact upstream revision, source
tree, lockfile, declared patch, retained license/notice, executable hash and full
native prompt. A workspace setting cannot replace these machine-scoped inputs.
Runtime inputs must be owned by the current user or root and must not be writable
by group or others. Use a dedicated prepared bundle (executables `0700` or `0555`,
report and prompt `0400` or `0444`), rather than relaxing checks for a group-writable
source checkout. Keep the original pinned source and its notices unchanged.
The core socket must belong to the current user with mode `0600` in an owned
`0700` directory. Existing runtime/model installation remains the operator's
explicit action. System Python 3 and bubblewrap must already be available.

Open a trusted local project and invoke the command. In a multi-folder workspace,
choose one folder; the other folders are not implicitly included. Enter the task
and confirm the selected read/write scope. Opening the extension or workspace
alone starts no runtime, model or network participation.

## Execution and privacy boundaries

The selected project is mounted as `/workspace` inside a disposable Linux
sandbox. The sandbox has a separate network/PID/mount namespace, no host user
home or inherited credentials, and only the prepared runtimes, system runtime
files, exact private core socket and selected project. It does not change host
DNS, routes or firewall. Broad system directories and overlap with launcher
inputs are refused as projects. The core processes private input locally; no
public cache, training, remote peer or OpenAI fallback is enabled by this slice.

The native runtime's workspace sandbox and approval policy remain in force.
When it requests command approval, the editor shows the exact command and its
relative working directory. **Run once** grants only that request, not a session,
future rule, network access or wider filesystem permissions. Unsupported tool
approval kinds and privilege/network expansion are refused. Workspace content
and model output do not grant permissions.

The agent can modify real files in the selected folder. Changes are **not** rolled
back on cancellation or failure; inspect Source Control and run the relevant
project checks. Prefer a dedicated working branch. The frontend does not label
a completed native turn as a verified completed task. Generated output is shown
as plaintext rather than executable HTML or Markdown.

Cancellation is forwarded to the exact thread and turn, including cancellation
during turn admission. Closing the extension also closes its owned session.
The launcher waits for provider/runtime shutdown; forced stops or uncertain
cleanup remain errors, not successful completion. Runtime stderr, bearer secrets
and raw prompts are not exported as diagnostics. The frontend does not retain
a durable conversation; the final untitled text can be saved only by the user.

## Verification scope and remaining work

Focused tests cover actual frontend/controller logic with synthetic protocol
events: explicit launch, user-only settings, scope selection, one-shot approval,
early native notifications, cancellation, EOF, cleanup failure and no automatic
startup. Launcher tests separately exercise process and namespace construction.
They do not stand in for the outstanding native editor/model trial.

A separate local protocol probe has passed through the production launcher and
the actual source-built app-server: initialize, open an ephemeral VOLPAROSSA
thread, unsubscribe, EOF and confirmed runtime/provider shutdown. Its core socket
was **synthetic and capability-only**: no turn, inference or tool execution was
requested, and VS Code/VSCodium itself was not launched. The temporary selected
project and read-only runtime staging were removed; original runtime bytes/modes
and host network state remained unchanged. Earlier attempts stopped before the
protocol handshake: first on group-writable source inputs, then because the
launcher rejected bubblewrap's own `PWD=/workspace`. Dedicated private staging
and an exact namespace-local PWD check resolved those launch blockers without
relaxing input ownership or importing host environment settings.

A separate **actual VSCodium UI** probe also passed: F1/Command Palette, the
capabilities view, the native-task input box, the consent dialog and clicking
Cancel before runtime startup. There were zero automatic requests and exactly
one capability request to a synthetic capability-only service. The editor exited
normally; its private profile was removed and host network state was unchanged.
This is UI admission evidence, **not** native inference or coding evidence:
**Run once**, model-driven edits and the final task-result view remain unproven.

## Disposable guest UI trial

`scripts/smoke_editor_ui.cjs` drives the real editor UI; it does **not** start an
editor, model, core service or VM. Execution requires `--execute --yes`, Linux
hostname `volparossa-alpha`, user `vpci` and KVM virtualization. Do not run the
model trial on the development host or bypass these guards.

The supervising guest launcher must first provide:

- The source-verified native runtime, complete upstream prompt and notices,
hash-verified Node 24 and VSCodium; reuse these assets, without downloading at
launch. VSCodium **1.135.06055**, commit
`1a46a584725d5dd330e0bcd7f5510f24990efcf2`, has actually opened a headless
workbench with `--ozone-platform=headless`; this version needs no Xvfb.
- A **real** owner-private `qwen3-0.6b-v1` conversation service with verified model
and Python-runtime provenance: two threads, 600 seconds per request, a 5 GiB
memory cgroup, no swap and a 2,700-second service window. Preserve the worker's
existing RSS/admission limits; an out-of-memory or admission failure is not
permission to weaken them.
- An isolated network/PID/mount/IPC environment, empty account home, no host
`DISPLAY`, Wayland, D-Bus or other host IPC mounts, and no inherited credentials.
Use an ordinary unprivileged user; **never add `--no-sandbox`**. Keep the core
socket and its parent at `0600`/`0700` and loopback CDP inside this environment.

The examples below assume that environment exposes this extension at `/extension`,
Node at `/opt/node`, and new owner-only directories under `/trial`. The project
must already be empty and mode `0700`, with a name such as `editor-ui-project-01`.
The reports directory must also be `0700`; output files must not already exist.

```sh
/opt/node /extension/scripts/smoke_editor_ui.cjs \
--prepare-project --execute --yes \
--project /trial/editor-ui-project-01 --output /trial/reports/prepare.json
```

Merge the explicit runtime/socket settings from the setup example into the
isolated profile's `User/settings.json`, alongside:

```json
{
"window.dialogStyle": "custom",
"workbench.startupEditor": "none",
"telemetry.telemetryLevel": "off",
"update.mode": "none",
"extensions.autoCheckUpdates": false,
"extensions.autoUpdate": false,
"security.workspace.trust.enabled": false
}
```

The last setting is **only for this disposable, explicitly selected fixture**,
not a recommended user default. Custom dialogs and English UI are required for
the real DOM selectors. Start the prepared editor inside the same isolation:

```sh
/usr/share/codium/codium --new-window --ozone-platform=headless --disable-gpu \
--disable-updates --disable-telemetry --disable-crash-reporter --locale=en \
--user-data-dir=/trial/profile --extensions-dir=/trial/extensions \
--extensionDevelopmentPath=/extension --skip-welcome --skip-release-notes \
--remote-debugging-address=127.0.0.1 --remote-debugging-port=9222 \
/trial/editor-ui-project-01
```

Once its workbench is ready, run from a separate supervised process:

```sh
/opt/node /extension/scripts/smoke_editor_ui.cjs --execute --yes \
--cdp http://127.0.0.1:9222 --project /trial/editor-ui-project-01 \
--output /trial/reports/ui.json --timeout-seconds 2400
```

The driver requires the unchanged prepared fixture, enters a real task, clicks
consent and approves only the bounded fixture commands. The model supplies the
edit expression. Success requires actual recorded read/edit/test actions,
unchanged helper code, a changed source file, an independent passing test and the
UI result displayed after runtime cleanup. The receipt contains closed statuses,
counts and hashes, not prompts, commands or private paths. **The full real-model
UI trial has not yet passed.** The parent supervisor still owns editor/core/VM
shutdown, private-profile/project removal and unchanged-host verification;
`ui.json` does not claim that broader cleanup.

The current prepared model is small; usable general coding quality is still to
be measured. Reviewable native diffs, durable multi-turn sessions, additional
tool types, broader platform support and eligible cooperative delegation remain
separate unfinished work. The core must own delegation and its privacy decision;
this local command must not silently publish a private project to peers.

Protocol reference: [official Codex app-server documentation](https://learn.chatgpt.com/docs/app-server).
6 changes: 5 additions & 1 deletion package.json
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@
"contributes": {
"commands": [
{"command": "volparossaCode.reviewSelection", "title": "VOLPAROSSA: Ask About Selected Code (Private, Local)"},
{"command": "volparossaCode.codingTask", "title": "VOLPAROSSA: Run Native Coding Task (Private, Local)"},
{"command": "volparossaCode.capabilities", "title": "VOLPAROSSA: Show Compute Capabilities"}
],
"configuration": {
Expand All @@ -25,10 +26,13 @@
"volparossaCode.privateSocket": {
"type": "string", "default": "", "scope": "machine",
"description": "Absolute same-owner Unix socket of an explicitly started VOLPAROSSA private-serve service. No service is started or downloaded by the extension."
},
"volparossaCode.nativeRuntime": {
"type": "object", "default": {}, "scope": "machine",
"description": "Explicit prepared native coding runtime inputs; see docs/NATIVE_EDITOR.md. User settings only. No runtime or model is downloaded, and opening a workspace starts nothing."
}
}
}
},
"scripts": {"test": "node --test tests/*.test.cjs", "check": "node --check src/extension.cjs && node --check src/app-server.cjs && node --check src/private-compute.cjs"}
}

Loading
Loading