Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 56 additions & 12 deletions docs/OPENCODE.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,12 @@ admission. Model quality and useful execution still require real-model evidence.

Before starting OpenCode, the owner validates the complete core capability reply,
including template, limits, private-local scope and negotiated `greedy_v1` policy.
Current native sessions also require `execution_error_version: 1`; an older core
without that capability is refused before starting the runtime. Both native smoke
profiles pin core `6a517b576baa17e7329661ee0476d1848081d114` for this contract,
including the preserved main integrations and reviewed source-quality fixes.
Legacy Q&A and conversation clients that do not opt in remain unchanged.

That one model identity binds the native catalog, all agent roles, provider,
task/session checks and trial receipt. Each provider request checks the core again;
a different profile or incompatible result is refused, not silently substituted.
Expand Down Expand Up @@ -112,13 +118,17 @@ for core cleanup before returning SDK-compatible SSE or JSON; this is not
token-by-token model streaming. Exhaustion remains `length`, not successful `stop`.
Cleanup-confirmed invalid/truncated output or an unmet tool choice returns a
terminal HTTP 422, so the pinned SDK and OpenCode do not blindly regenerate that
unusable turn. Busy/execution/transport availability and uncertain cleanup remain
separate failures; this change does not make an incomplete answer usable.
unusable turn. The negotiated `execution_budget_exceeded` response also maps to
422, but only for the matching admitted task after confirmed core cleanup. It
does not claim a completed model response. Busy, other execution/transport
failures and uncertain cleanup remain separate; this change does not make an
incomplete answer usable.
An already reported, known terminal task failure is separate from runtime cleanup:
the task still fails, while a confirmed session/provider/process shutdown can
succeed. Unknown/protocol errors and any unconfirmed cleanup still fail closed.
The provider also counts cleanup for a correlated, admitted task ending with
the core's terminal `execution_failed` or `cancelled` response, or a valid result
the core's terminal `execution_failed`, `cancelled` or negotiated
`execution_budget_exceeded` response, or a valid result
that races local cancellation. These remain failed requests, not model results.
The receipt is retained per rejection inside the checked transport; an error code
alone, admission alone, a cancellation acknowledgement or a disconnect cannot
Expand Down Expand Up @@ -458,9 +468,31 @@ missing/conflicting result evidence is an error, not a sampled fallback.
An older ordinary conversation client can still use its unchanged handshake;
omitting the policy retains the core's previous profile behavior.

This corrects a real contract mismatch: the earlier adapter accepted zero
temperature while the Qwen worker used its default sampled 0.7/0.8/20 profile.
The fixed worker passes `do_sample:false,num_beams:1`; model, token budgets,
An actual-runtime regression exposed an additional configuration mismatch:
pinned OpenCode omits temperature unless the custom model advertises that
capability. Setting only `agent.temperature:0` did not select greedy execution.
Both supported profiles now explicitly set model `temperature:true`, retaining
the existing zero setting on every coding role. Prompts, model identities,
permissions and resource budgets are unchanged; this corrects which generation
policy actually reaches the core, rather than proving better model quality.

The pinned native runtime has now been exercised with the production launcher
and provider but **synthetic** core/model replies. Each of `invalid_output`,
`wire_truncated` and `execution_budget_exceeded` made exactly one coding request,
with observed `greedy_v1`, no native retry, no approval and confirmed session
cleanup; the disposable project remained unchanged. The terminal budget error
was not counted as a completed or incomplete model result. The synthetic result
fixture mirrors only the policy actually present on the request; negotiated
support alone is not a selected policy. This proves runtime policy forwarding
and terminal-error behavior, not actual inference or private peer execution.
Local report `native-errors-budget-greedy-df481-20261004.json` was captured on
the explicitly modified `df481f3f` candidate; SHA-256:
`a08f26ab2c61b2653949dadf06f45b9402066383ee651d573a5b14f9623e3786`.
Earlier real-model failures retain their original sources and outcomes.

An earlier adapter correction addressed a separate contract mismatch: it accepted
zero temperature while the Qwen worker used its sampled 0.7/0.8/20 profile.
The worker's greedy path passes `do_sample:false,num_beams:1`; model, token budgets,
tool permissions, task and independent success checks remain unchanged.
Protocol/backend-double checks verify the wiring, not model quality. The earlier
[run `37066003771`](https://github.com/VOLPAROSSA/volparossa-code/actions/runs/37066003771)
Expand All @@ -473,12 +505,24 @@ stop, and greedy generation does not guarantee a completed coding task.
`scripts/smoke_opencode_inference.py` prepares one explicit disposable Debian 13
KVM trial; it does not install or run the model on the development host. Its
`pack` mode captures each Code source/runtime file by hash and the exact core
archive selected by its explicit model profile. The default Qwen3-0.6B profile
retains core `845cc84d0d0b766ab1c5227231dbf6c8eaeb8cc3`; the separate
`qwen3-4b-instruct-2507-v1` candidate binds core
`1297f8f1a5d163d802efd066c51a950b95588fa5`. This includes the bounded
provisioning-timeout recovery from core `3aa0e2d0` and existing closed worker
diagnostics. The previous `37209881216` attempt remains failed before model
archive selected by its explicit model profile. Both current profiles pin core
`6a517b576baa17e7329661ee0476d1848081d114`, while retaining their different
models and resource profiles. Historical 0.6B trials used `845cc84d`; the last
4B trial used `1297f8f1a5d163d802efd066c51a950b95588fa5`, including the bounded
provisioning-timeout recovery from `3aa0e2d0`. Neither those historical outcomes
nor the synthetic native-runtime probe proves the current real-model candidate.

The earlier `a5246942` candidate's
[Quality run `37225595367`](https://github.com/VOLPAROSSA/volparossa/actions/runs/37225595367)
remains **failed**: strict Clippy rejected two 103-line diagnostic functions.
No native model trial was dispatched for that pair. The follow-up extracts the
existing snapshot parser and moves byte-identical Python test assertions into a
constant, without changing validation, model behavior, resources or acceptance
criteria. Its source checks are pending; this is not a rerun or a passing
real-model result. Original job `111504476222` log SHA-256:
`1ef3aca56b4052b9b56e2f9a469543e07cfdb1718d5a8706750c286f40f6d41c`.

The previous `37209881216` attempt remains failed before model
execution. [Trial `37217032474`](https://github.com/VOLPAROSSA/volparossa-code/actions/runs/37217032474)
on the previous core `f25352df` completed the pinned 4B provisioning and returned one actual model result,
but the native task failed at `invalid_output` after 152,995 ms. No tools, edits
Expand Down
4 changes: 2 additions & 2 deletions scripts/opencode_session.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -135,11 +135,11 @@ async function runSession({input, output, events = process}, hooks = {}) {
for (const name of ['SIGTERM', 'SIGINT', 'SIGHUP']) events.once(name, broken);
try {
const caps = capabilities(await (hooks.preflight ?? (async () => {
const core = new PrivateConversation('/opt/core/compute.sock', {generationPolicyVersion: 1});
const core = new PrivateConversation('/opt/core/compute.sock', {generationPolicyVersion: 1, executionErrorVersion: 1});
try {
return await core.connect();
} finally { core.close(); }
}))(), 1);
}))(), 1, 1);
if (!isCodingModel(caps.model_profile) || !caps.native_tool_template || !caps.local_only ||
caps.quarantined || !caps.generation_policies.includes('greedy_v1')) throw Error('capabilities');
// The owner's already selected core determines this session's one model.
Expand Down
6 changes: 3 additions & 3 deletions scripts/smoke_opencode_inference.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,12 +27,12 @@
import uuid

ROOT = Path(__file__).resolve().parents[1]
CORE = '845cc84d0d0b766ab1c5227231dbf6c8eaeb8cc3'
CORE = '6a517b576baa17e7329661ee0476d1848081d114'
MODEL = 'qwen3-0.6b-v1'
GIB = 1024 ** 3
LARGE_MODEL = 'qwen3-4b-instruct-2507-v1'
# Separate reviewed source; the default profile retains its original core pin.
LARGE_CORE = '1297f8f1a5d163d802efd066c51a950b95588fa5'
# Both profiles require the reviewed terminal-budget IPC; model and budgets stay distinct.
LARGE_CORE = '6a517b576baa17e7329661ee0476d1848081d114'
MODEL_PROFILES = (MODEL, LARGE_MODEL)
ENV = {'PATH': '/usr/bin:/bin', 'LANG': 'C.UTF-8', 'PYTHONDONTWRITEBYTECODE': '1'}
IMAGE_NAME = 'debian-13-genericcloud-amd64-20260826-2582.qcow2'
Expand Down
9 changes: 6 additions & 3 deletions src/chat-completions-provider.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -241,7 +241,7 @@ async function startChatCompletionsProvider({ socketPath, model, diagnostics = f
if (closing || active) { errorReply(response, 503, 'busy'); return; }
if (Number(request.headers['content-length'] ?? 0) > HTTP_BYTES) { errorReply(response, 413, 'request_bound'); return; }
const controller = new AbortController();
const client = new PrivateConversation(socketPath, { generationPolicyVersion: 1 });
const client = new PrivateConversation(socketPath, { generationPolicyVersion: 1, executionErrorVersion: 1 });
const owner = { controller, client, done: null };
active = owner;
response.once('close', () => { if (!response.writableFinished) controller.abort(); });
Expand Down Expand Up @@ -287,11 +287,14 @@ async function startChatCompletionsProvider({ socketPath, model, diagnostics = f
}
const code = error.message?.startsWith('private_compute_') ? error.code : 'provider_failed';
if (diagnostics) count(summary.request_errors, PROVIDER_ERRORS.includes(code) ? code : 'other');
// These two errors follow a terminal, cleanup-confirmed model result.
// Output errors follow a terminal, cleanup-confirmed model result. The
// negotiated execution budget error instead follows confirmed cleanup
// without a model result; it must not restart the same expensive turn.
// OpenCode v1.18.34 retries every 5xx, so 502 would repeatedly regenerate
// the same unusable greedy turn. Preserve transient/uncertain failures.
const status = ['busy', 'cleanup_unconfirmed', 'socket_unavailable', 'execution_failed'].includes(code) ? 503 :
['invalid_model_output', 'tool_choice_not_met'].includes(code) ? 422 : code === 'request_bound' ? 413 : 400;
['invalid_model_output', 'tool_choice_not_met', 'execution_budget_exceeded'].includes(code) ? 422 :
code === 'request_bound' ? 413 : 400;
errorReply(response, status, code);
} finally {
client.close();
Expand Down
7 changes: 6 additions & 1 deletion src/opencode-bridge.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ const FAILURE_CODES = new Set([
'event_disposed', 'event_response', 'event_bound', 'event_invalid', 'event_closed',
'event_transport', 'session', 'model', 'prompt'].map(code => `opencode_${code}`),
]);
const PROVIDER_ERRORS = Object.freeze(['execution_failed', 'invalid_model_output', 'tool_choice_not_met',
const PROVIDER_ERRORS = Object.freeze(['execution_failed', 'execution_budget_exceeded', 'invalid_model_output', 'tool_choice_not_met',
'invalid_conversation', 'request_bound', 'busy', 'cancelled', 'cleanup_unconfirmed',
'socket_unavailable', 'socket_error', 'disconnected', 'invalid_response',
'incompatible_capabilities', 'provider_failed', 'other']);
Expand Down Expand Up @@ -63,6 +63,11 @@ function validTaskDiagnostic(value) {
}
function validProviderDiagnostic(value) {
const template = emptyProviderDiagnostic();
// Historical v1 receipts predate the negotiated budget code. Accept either
// exact closed schema; never fabricate the missing counter or relax extras.
if (record(value?.request_errors) && !Object.hasOwn(value.request_errors, 'execution_budget_exceeded')) {
delete template.request_errors.execution_budget_exceeded;
}
const matches = (actual, expected) => record(actual) && Object.keys(actual).length === Object.keys(expected).length
&& Object.entries(expected).every(([key, item]) => Object.hasOwn(actual, key) && (record(item)
? matches(actual[key], item) : typeof item === 'number'
Expand Down
5 changes: 4 additions & 1 deletion src/opencode-config.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,10 @@ function runtimeSettings({baseUrl, bearerToken, password, cooperative = false, m
models: {[model]: {name: 'VOLPAROSSA core conversation',
limit: {context: expectedLimits(model).model_context_tokens, output: expectedLimits(model).max_new_tokens},
tool_call: true, reasoning: false,
modalities: {input: ['text'], output: ['text']}}},
modalities: {input: ['text'], output: ['text']},
// Pinned OpenCode otherwise drops agent.temperature, silently omitting
// the zero-temperature request that selects core's negotiated greedy_v1.
temperature: true}},
}},
};
const env = {
Expand Down
9 changes: 7 additions & 2 deletions src/private-compute.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -275,13 +275,18 @@ class PrivateCompute {
}
}

// Only the negotiated conversation subclass can recognize this additive code.
// Q&A and legacy conversations retain their original closed error vocabulary.
_allowsExecutionBudgetError(_message) { return false; }

_response(message) {
requireValue(object(message) && message.version === 1 && typeof message.event === 'string');
requireValue((typeof message.id === 'string' && /^[0-9a-f]{32}$/.test(message.id)) ||
(message.id === null && message.event === 'error' && message.code === 'invalid_request'));
if (message.event === 'error') {
keys(message, ['version', 'id', 'event', 'code']);
requireValue(REMOTE_ERRORS.has(message.code));
const budgetError = message.code === 'execution_budget_exceeded' && this._allowsExecutionBudgetError(message);
requireValue(REMOTE_ERRORS.has(message.code) || budgetError);
const error = failure(message.code);
if (message.id === null) throw error;
requireValue(message.id === this.handshake?.id || message.id === this.pending?.id ||
Expand All @@ -298,7 +303,7 @@ class PrivateCompute {
// through cleanup. Uncertainty takes precedence as cleanup_unconfirmed.
// Do not infer this receipt from an error code before admission or from a
// transport failure that merely has the same message.
if (this.pending.admitted && ['execution_failed', 'cancelled'].includes(message.code)) {
if (this.pending.admitted && (['execution_failed', 'cancelled'].includes(message.code) || budgetError)) {
cleanedFailures.add(error);
}
this._settle(error);
Expand Down
20 changes: 15 additions & 5 deletions src/private-conversation.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -63,11 +63,12 @@ function equalLimits(value, expected, optional = []) {
keys(value, Object.keys(expected), optional);
check(Object.entries(expected).every(([key, item]) => value[key] === item), 'incompatible_capabilities');
}
function capabilities(value, generationPolicyVersion) {
function capabilities(value, generationPolicyVersion, executionErrorVersion) {
check(object(value));
const generation = generationPolicyVersion === 1 ? ['generation_policy_version', 'generation_policies'] : [];
const execution = executionErrorVersion === 1 ? ['execution_error_version'] : [];
equalLimits(value, { ...expectedLimits(value.model_profile), execution_slots: 1,
max_request_bytes: requestLimit(value.model_profile), max_response_bytes: 65536 }, ['max_seconds', 'quarantined', ...generation]);
max_request_bytes: requestLimit(value.model_profile), max_response_bytes: 65536 }, ['max_seconds', 'quarantined', ...generation, ...execution]);
check(Number.isInteger(value.max_seconds) && value.max_seconds >= 1 && value.max_seconds <= 600 &&
typeof value.quarantined === 'boolean', 'incompatible_capabilities');
if (generationPolicyVersion === 1) {
Expand All @@ -76,6 +77,7 @@ function capabilities(value, generationPolicyVersion) {
JSON.stringify(value.generation_policies) === JSON.stringify(expected), 'unsupported_generation_policy');
Object.freeze(value.generation_policies);
}
if (executionErrorVersion === 1) check(value.execution_error_version === 1, 'incompatible_capabilities');
return Object.freeze(value);
}

Expand Down Expand Up @@ -168,18 +170,21 @@ function validateResult(value, caps, input) {
}

class PrivateConversation extends PrivateCompute {
constructor(socketPath, { generationPolicyVersion } = {}) {
constructor(socketPath, { generationPolicyVersion, executionErrorVersion } = {}) {
super(socketPath);
check(generationPolicyVersion === undefined || generationPolicyVersion === 1, 'unsupported_generation_policy');
this.generationPolicyVersion = generationPolicyVersion;
check(executionErrorVersion === undefined || executionErrorVersion === 1, 'unsupported_execution_error_version');
this.executionErrorVersion = executionErrorVersion;
}
// These hooks reuse only the already-tested transport. Q&A callers and bytes
// remain unchanged; this instance cannot silently use the old handshake.
_send(id, operation) {
// The separate conversation family may advertise a larger request envelope;
// the legacy Q&A class and its 32KiB framing remain byte-for-byte unchanged.
if (operation.type === 'capabilities') operation = { type: 'conversation_capabilities',
...(this.generationPolicyVersion === 1 ? { generation_policy_version: 1 } : {}) };
...(this.generationPolicyVersion === 1 ? { generation_policy_version: 1 } : {}),
...(this.executionErrorVersion === 1 ? { execution_error_version: 1 } : {}) };
try {
const body = Buffer.from(JSON.stringify({ version: 1, id, operation }));
check(body.length > 0 && body.length <= (this.caps?.max_request_bytes ?? 32768), 'request_bound');
Expand All @@ -203,18 +208,23 @@ class PrivateConversation extends PrivateCompute {
return new Promise((resolve, reject) => {
const abort = () => this._cancel();
this.pending = { id, input, resolve, reject, signal, abort, admitted: false, cancelled: false,
executionErrorVersion: this.caps.execution_error_version === 1 ? 1 : undefined,
timer: setTimeout(() => this._cancel(), this.caps.max_seconds * 1000 + 5000) };
signal?.addEventListener('abort', abort, { once: true });
this._send(id, { type: 'submit_conversation', conversation: input });
if (signal?.aborted) abort();
});
}
_allowsExecutionBudgetError(message) {
return this.state === 'open' && this.pending?.admitted === true &&
this.pending.executionErrorVersion === 1 && message.id === this.pending.id && !this.cancels.has(message.id);
}
_response(message) {
check(object(message) && message.version === 1 && message.event !== 'capabilities', 'invalid_response');
if (message.event === 'conversation_capabilities') {
keys(message, ['version', 'id', 'event', 'capabilities']);
check(this.state === 'connecting' && message.id === this.handshake?.id, 'invalid_response');
this.caps = capabilities(message.capabilities, this.generationPolicyVersion);
this.caps = capabilities(message.capabilities, this.generationPolicyVersion, this.executionErrorVersion);
this.state = 'open';
clearTimeout(this.handshake.timer);
this.handshake.resolve(this.caps);
Expand Down
Loading
Loading