diff --git a/docs/OPENCODE.md b/docs/OPENCODE.md index a9031d6..95b7547 100644 --- a/docs/OPENCODE.md +++ b/docs/OPENCODE.md @@ -75,6 +75,12 @@ admission. Model quality and useful execution still require real-model evidence. Before starting OpenCode, the owner validates the complete core capability reply, including template, limits, private-local scope and negotiated `greedy_v1` policy. +Current native sessions also require `execution_error_version: 1`; an older core +without that capability is refused before starting the runtime. Both native smoke +profiles pin core `6a517b576baa17e7329661ee0476d1848081d114` for this contract, +including the preserved main integrations and reviewed source-quality fixes. +Legacy Q&A and conversation clients that do not opt in remain unchanged. + That one model identity binds the native catalog, all agent roles, provider, task/session checks and trial receipt. Each provider request checks the core again; a different profile or incompatible result is refused, not silently substituted. @@ -112,13 +118,17 @@ for core cleanup before returning SDK-compatible SSE or JSON; this is not token-by-token model streaming. Exhaustion remains `length`, not successful `stop`. Cleanup-confirmed invalid/truncated output or an unmet tool choice returns a terminal HTTP 422, so the pinned SDK and OpenCode do not blindly regenerate that -unusable turn. Busy/execution/transport availability and uncertain cleanup remain -separate failures; this change does not make an incomplete answer usable. +unusable turn. The negotiated `execution_budget_exceeded` response also maps to +422, but only for the matching admitted task after confirmed core cleanup. It +does not claim a completed model response. Busy, other execution/transport +failures and uncertain cleanup remain separate; this change does not make an +incomplete answer usable. An already reported, known terminal task failure is separate from runtime cleanup: the task still fails, while a confirmed session/provider/process shutdown can succeed. Unknown/protocol errors and any unconfirmed cleanup still fail closed. The provider also counts cleanup for a correlated, admitted task ending with -the core's terminal `execution_failed` or `cancelled` response, or a valid result +the core's terminal `execution_failed`, `cancelled` or negotiated +`execution_budget_exceeded` response, or a valid result that races local cancellation. These remain failed requests, not model results. The receipt is retained per rejection inside the checked transport; an error code alone, admission alone, a cancellation acknowledgement or a disconnect cannot @@ -458,9 +468,31 @@ missing/conflicting result evidence is an error, not a sampled fallback. An older ordinary conversation client can still use its unchanged handshake; omitting the policy retains the core's previous profile behavior. -This corrects a real contract mismatch: the earlier adapter accepted zero -temperature while the Qwen worker used its default sampled 0.7/0.8/20 profile. -The fixed worker passes `do_sample:false,num_beams:1`; model, token budgets, +An actual-runtime regression exposed an additional configuration mismatch: +pinned OpenCode omits temperature unless the custom model advertises that +capability. Setting only `agent.temperature:0` did not select greedy execution. +Both supported profiles now explicitly set model `temperature:true`, retaining +the existing zero setting on every coding role. Prompts, model identities, +permissions and resource budgets are unchanged; this corrects which generation +policy actually reaches the core, rather than proving better model quality. + +The pinned native runtime has now been exercised with the production launcher +and provider but **synthetic** core/model replies. Each of `invalid_output`, +`wire_truncated` and `execution_budget_exceeded` made exactly one coding request, +with observed `greedy_v1`, no native retry, no approval and confirmed session +cleanup; the disposable project remained unchanged. The terminal budget error +was not counted as a completed or incomplete model result. The synthetic result +fixture mirrors only the policy actually present on the request; negotiated +support alone is not a selected policy. This proves runtime policy forwarding +and terminal-error behavior, not actual inference or private peer execution. +Local report `native-errors-budget-greedy-df481-20261004.json` was captured on +the explicitly modified `df481f3f` candidate; SHA-256: +`a08f26ab2c61b2653949dadf06f45b9402066383ee651d573a5b14f9623e3786`. +Earlier real-model failures retain their original sources and outcomes. + +An earlier adapter correction addressed a separate contract mismatch: it accepted +zero temperature while the Qwen worker used its sampled 0.7/0.8/20 profile. +The worker's greedy path passes `do_sample:false,num_beams:1`; model, token budgets, tool permissions, task and independent success checks remain unchanged. Protocol/backend-double checks verify the wiring, not model quality. The earlier [run `37066003771`](https://github.com/VOLPAROSSA/volparossa-code/actions/runs/37066003771) @@ -473,12 +505,24 @@ stop, and greedy generation does not guarantee a completed coding task. `scripts/smoke_opencode_inference.py` prepares one explicit disposable Debian 13 KVM trial; it does not install or run the model on the development host. Its `pack` mode captures each Code source/runtime file by hash and the exact core -archive selected by its explicit model profile. The default Qwen3-0.6B profile -retains core `845cc84d0d0b766ab1c5227231dbf6c8eaeb8cc3`; the separate -`qwen3-4b-instruct-2507-v1` candidate binds core -`1297f8f1a5d163d802efd066c51a950b95588fa5`. This includes the bounded -provisioning-timeout recovery from core `3aa0e2d0` and existing closed worker -diagnostics. The previous `37209881216` attempt remains failed before model +archive selected by its explicit model profile. Both current profiles pin core +`6a517b576baa17e7329661ee0476d1848081d114`, while retaining their different +models and resource profiles. Historical 0.6B trials used `845cc84d`; the last +4B trial used `1297f8f1a5d163d802efd066c51a950b95588fa5`, including the bounded +provisioning-timeout recovery from `3aa0e2d0`. Neither those historical outcomes +nor the synthetic native-runtime probe proves the current real-model candidate. + +The earlier `a5246942` candidate's +[Quality run `37225595367`](https://github.com/VOLPAROSSA/volparossa/actions/runs/37225595367) +remains **failed**: strict Clippy rejected two 103-line diagnostic functions. +No native model trial was dispatched for that pair. The follow-up extracts the +existing snapshot parser and moves byte-identical Python test assertions into a +constant, without changing validation, model behavior, resources or acceptance +criteria. Its source checks are pending; this is not a rerun or a passing +real-model result. Original job `111504476222` log SHA-256: +`1ef3aca56b4052b9b56e2f9a469543e07cfdb1718d5a8706750c286f40f6d41c`. + +The previous `37209881216` attempt remains failed before model execution. [Trial `37217032474`](https://github.com/VOLPAROSSA/volparossa-code/actions/runs/37217032474) on the previous core `f25352df` completed the pinned 4B provisioning and returned one actual model result, but the native task failed at `invalid_output` after 152,995 ms. No tools, edits diff --git a/scripts/opencode_session.cjs b/scripts/opencode_session.cjs index 1bb3fbe..35c974e 100644 --- a/scripts/opencode_session.cjs +++ b/scripts/opencode_session.cjs @@ -135,11 +135,11 @@ async function runSession({input, output, events = process}, hooks = {}) { for (const name of ['SIGTERM', 'SIGINT', 'SIGHUP']) events.once(name, broken); try { const caps = capabilities(await (hooks.preflight ?? (async () => { - const core = new PrivateConversation('/opt/core/compute.sock', {generationPolicyVersion: 1}); + const core = new PrivateConversation('/opt/core/compute.sock', {generationPolicyVersion: 1, executionErrorVersion: 1}); try { return await core.connect(); } finally { core.close(); } - }))(), 1); + }))(), 1, 1); if (!isCodingModel(caps.model_profile) || !caps.native_tool_template || !caps.local_only || caps.quarantined || !caps.generation_policies.includes('greedy_v1')) throw Error('capabilities'); // The owner's already selected core determines this session's one model. diff --git a/scripts/smoke_opencode_inference.py b/scripts/smoke_opencode_inference.py index d9ff175..e4adbfc 100644 --- a/scripts/smoke_opencode_inference.py +++ b/scripts/smoke_opencode_inference.py @@ -27,12 +27,12 @@ import uuid ROOT = Path(__file__).resolve().parents[1] -CORE = '845cc84d0d0b766ab1c5227231dbf6c8eaeb8cc3' +CORE = '6a517b576baa17e7329661ee0476d1848081d114' MODEL = 'qwen3-0.6b-v1' GIB = 1024 ** 3 LARGE_MODEL = 'qwen3-4b-instruct-2507-v1' -# Separate reviewed source; the default profile retains its original core pin. -LARGE_CORE = '1297f8f1a5d163d802efd066c51a950b95588fa5' +# Both profiles require the reviewed terminal-budget IPC; model and budgets stay distinct. +LARGE_CORE = '6a517b576baa17e7329661ee0476d1848081d114' MODEL_PROFILES = (MODEL, LARGE_MODEL) ENV = {'PATH': '/usr/bin:/bin', 'LANG': 'C.UTF-8', 'PYTHONDONTWRITEBYTECODE': '1'} IMAGE_NAME = 'debian-13-genericcloud-amd64-20260826-2582.qcow2' diff --git a/src/chat-completions-provider.cjs b/src/chat-completions-provider.cjs index 0095575..84fe361 100644 --- a/src/chat-completions-provider.cjs +++ b/src/chat-completions-provider.cjs @@ -241,7 +241,7 @@ async function startChatCompletionsProvider({ socketPath, model, diagnostics = f if (closing || active) { errorReply(response, 503, 'busy'); return; } if (Number(request.headers['content-length'] ?? 0) > HTTP_BYTES) { errorReply(response, 413, 'request_bound'); return; } const controller = new AbortController(); - const client = new PrivateConversation(socketPath, { generationPolicyVersion: 1 }); + const client = new PrivateConversation(socketPath, { generationPolicyVersion: 1, executionErrorVersion: 1 }); const owner = { controller, client, done: null }; active = owner; response.once('close', () => { if (!response.writableFinished) controller.abort(); }); @@ -287,11 +287,14 @@ async function startChatCompletionsProvider({ socketPath, model, diagnostics = f } const code = error.message?.startsWith('private_compute_') ? error.code : 'provider_failed'; if (diagnostics) count(summary.request_errors, PROVIDER_ERRORS.includes(code) ? code : 'other'); - // These two errors follow a terminal, cleanup-confirmed model result. + // Output errors follow a terminal, cleanup-confirmed model result. The + // negotiated execution budget error instead follows confirmed cleanup + // without a model result; it must not restart the same expensive turn. // OpenCode v1.18.34 retries every 5xx, so 502 would repeatedly regenerate // the same unusable greedy turn. Preserve transient/uncertain failures. const status = ['busy', 'cleanup_unconfirmed', 'socket_unavailable', 'execution_failed'].includes(code) ? 503 : - ['invalid_model_output', 'tool_choice_not_met'].includes(code) ? 422 : code === 'request_bound' ? 413 : 400; + ['invalid_model_output', 'tool_choice_not_met', 'execution_budget_exceeded'].includes(code) ? 422 : + code === 'request_bound' ? 413 : 400; errorReply(response, status, code); } finally { client.close(); diff --git a/src/opencode-bridge.cjs b/src/opencode-bridge.cjs index 1a743c4..7476012 100644 --- a/src/opencode-bridge.cjs +++ b/src/opencode-bridge.cjs @@ -17,7 +17,7 @@ const FAILURE_CODES = new Set([ 'event_disposed', 'event_response', 'event_bound', 'event_invalid', 'event_closed', 'event_transport', 'session', 'model', 'prompt'].map(code => `opencode_${code}`), ]); -const PROVIDER_ERRORS = Object.freeze(['execution_failed', 'invalid_model_output', 'tool_choice_not_met', +const PROVIDER_ERRORS = Object.freeze(['execution_failed', 'execution_budget_exceeded', 'invalid_model_output', 'tool_choice_not_met', 'invalid_conversation', 'request_bound', 'busy', 'cancelled', 'cleanup_unconfirmed', 'socket_unavailable', 'socket_error', 'disconnected', 'invalid_response', 'incompatible_capabilities', 'provider_failed', 'other']); @@ -63,6 +63,11 @@ function validTaskDiagnostic(value) { } function validProviderDiagnostic(value) { const template = emptyProviderDiagnostic(); + // Historical v1 receipts predate the negotiated budget code. Accept either + // exact closed schema; never fabricate the missing counter or relax extras. + if (record(value?.request_errors) && !Object.hasOwn(value.request_errors, 'execution_budget_exceeded')) { + delete template.request_errors.execution_budget_exceeded; + } const matches = (actual, expected) => record(actual) && Object.keys(actual).length === Object.keys(expected).length && Object.entries(expected).every(([key, item]) => Object.hasOwn(actual, key) && (record(item) ? matches(actual[key], item) : typeof item === 'number' diff --git a/src/opencode-config.cjs b/src/opencode-config.cjs index fc70607..2c2b484 100644 --- a/src/opencode-config.cjs +++ b/src/opencode-config.cjs @@ -51,7 +51,10 @@ function runtimeSettings({baseUrl, bearerToken, password, cooperative = false, m models: {[model]: {name: 'VOLPAROSSA core conversation', limit: {context: expectedLimits(model).model_context_tokens, output: expectedLimits(model).max_new_tokens}, tool_call: true, reasoning: false, - modalities: {input: ['text'], output: ['text']}}}, + modalities: {input: ['text'], output: ['text']}, + // Pinned OpenCode otherwise drops agent.temperature, silently omitting + // the zero-temperature request that selects core's negotiated greedy_v1. + temperature: true}}, }}, }; const env = { diff --git a/src/private-compute.cjs b/src/private-compute.cjs index 09f0d7d..06a9983 100644 --- a/src/private-compute.cjs +++ b/src/private-compute.cjs @@ -275,13 +275,18 @@ class PrivateCompute { } } + // Only the negotiated conversation subclass can recognize this additive code. + // Q&A and legacy conversations retain their original closed error vocabulary. + _allowsExecutionBudgetError(_message) { return false; } + _response(message) { requireValue(object(message) && message.version === 1 && typeof message.event === 'string'); requireValue((typeof message.id === 'string' && /^[0-9a-f]{32}$/.test(message.id)) || (message.id === null && message.event === 'error' && message.code === 'invalid_request')); if (message.event === 'error') { keys(message, ['version', 'id', 'event', 'code']); - requireValue(REMOTE_ERRORS.has(message.code)); + const budgetError = message.code === 'execution_budget_exceeded' && this._allowsExecutionBudgetError(message); + requireValue(REMOTE_ERRORS.has(message.code) || budgetError); const error = failure(message.code); if (message.id === null) throw error; requireValue(message.id === this.handshake?.id || message.id === this.pending?.id || @@ -298,7 +303,7 @@ class PrivateCompute { // through cleanup. Uncertainty takes precedence as cleanup_unconfirmed. // Do not infer this receipt from an error code before admission or from a // transport failure that merely has the same message. - if (this.pending.admitted && ['execution_failed', 'cancelled'].includes(message.code)) { + if (this.pending.admitted && (['execution_failed', 'cancelled'].includes(message.code) || budgetError)) { cleanedFailures.add(error); } this._settle(error); diff --git a/src/private-conversation.cjs b/src/private-conversation.cjs index 9585c6b..79663d8 100644 --- a/src/private-conversation.cjs +++ b/src/private-conversation.cjs @@ -63,11 +63,12 @@ function equalLimits(value, expected, optional = []) { keys(value, Object.keys(expected), optional); check(Object.entries(expected).every(([key, item]) => value[key] === item), 'incompatible_capabilities'); } -function capabilities(value, generationPolicyVersion) { +function capabilities(value, generationPolicyVersion, executionErrorVersion) { check(object(value)); const generation = generationPolicyVersion === 1 ? ['generation_policy_version', 'generation_policies'] : []; + const execution = executionErrorVersion === 1 ? ['execution_error_version'] : []; equalLimits(value, { ...expectedLimits(value.model_profile), execution_slots: 1, - max_request_bytes: requestLimit(value.model_profile), max_response_bytes: 65536 }, ['max_seconds', 'quarantined', ...generation]); + max_request_bytes: requestLimit(value.model_profile), max_response_bytes: 65536 }, ['max_seconds', 'quarantined', ...generation, ...execution]); check(Number.isInteger(value.max_seconds) && value.max_seconds >= 1 && value.max_seconds <= 600 && typeof value.quarantined === 'boolean', 'incompatible_capabilities'); if (generationPolicyVersion === 1) { @@ -76,6 +77,7 @@ function capabilities(value, generationPolicyVersion) { JSON.stringify(value.generation_policies) === JSON.stringify(expected), 'unsupported_generation_policy'); Object.freeze(value.generation_policies); } + if (executionErrorVersion === 1) check(value.execution_error_version === 1, 'incompatible_capabilities'); return Object.freeze(value); } @@ -168,10 +170,12 @@ function validateResult(value, caps, input) { } class PrivateConversation extends PrivateCompute { - constructor(socketPath, { generationPolicyVersion } = {}) { + constructor(socketPath, { generationPolicyVersion, executionErrorVersion } = {}) { super(socketPath); check(generationPolicyVersion === undefined || generationPolicyVersion === 1, 'unsupported_generation_policy'); this.generationPolicyVersion = generationPolicyVersion; + check(executionErrorVersion === undefined || executionErrorVersion === 1, 'unsupported_execution_error_version'); + this.executionErrorVersion = executionErrorVersion; } // These hooks reuse only the already-tested transport. Q&A callers and bytes // remain unchanged; this instance cannot silently use the old handshake. @@ -179,7 +183,8 @@ class PrivateConversation extends PrivateCompute { // The separate conversation family may advertise a larger request envelope; // the legacy Q&A class and its 32KiB framing remain byte-for-byte unchanged. if (operation.type === 'capabilities') operation = { type: 'conversation_capabilities', - ...(this.generationPolicyVersion === 1 ? { generation_policy_version: 1 } : {}) }; + ...(this.generationPolicyVersion === 1 ? { generation_policy_version: 1 } : {}), + ...(this.executionErrorVersion === 1 ? { execution_error_version: 1 } : {}) }; try { const body = Buffer.from(JSON.stringify({ version: 1, id, operation })); check(body.length > 0 && body.length <= (this.caps?.max_request_bytes ?? 32768), 'request_bound'); @@ -203,18 +208,23 @@ class PrivateConversation extends PrivateCompute { return new Promise((resolve, reject) => { const abort = () => this._cancel(); this.pending = { id, input, resolve, reject, signal, abort, admitted: false, cancelled: false, + executionErrorVersion: this.caps.execution_error_version === 1 ? 1 : undefined, timer: setTimeout(() => this._cancel(), this.caps.max_seconds * 1000 + 5000) }; signal?.addEventListener('abort', abort, { once: true }); this._send(id, { type: 'submit_conversation', conversation: input }); if (signal?.aborted) abort(); }); } + _allowsExecutionBudgetError(message) { + return this.state === 'open' && this.pending?.admitted === true && + this.pending.executionErrorVersion === 1 && message.id === this.pending.id && !this.cancels.has(message.id); + } _response(message) { check(object(message) && message.version === 1 && message.event !== 'capabilities', 'invalid_response'); if (message.event === 'conversation_capabilities') { keys(message, ['version', 'id', 'event', 'capabilities']); check(this.state === 'connecting' && message.id === this.handshake?.id, 'invalid_response'); - this.caps = capabilities(message.capabilities, this.generationPolicyVersion); + this.caps = capabilities(message.capabilities, this.generationPolicyVersion, this.executionErrorVersion); this.state = 'open'; clearTimeout(this.handshake.timer); this.handshake.resolve(this.caps); diff --git a/tests/chat-completions-provider.test.cjs b/tests/chat-completions-provider.test.cjs index 33f4980..698ab6b 100644 --- a/tests/chat-completions-provider.test.cjs +++ b/tests/chat-completions-provider.test.cjs @@ -11,7 +11,8 @@ const { startChatCompletionsProvider, toConversation } = require('../src/chat-co const { caps: legacyCaps, result: legacyResult, reply, fixture } = require('./conversation-fixture.cjs'); // These are synthetic replies for the explicitly negotiated worker policy. -const caps = model => ({ ...legacyCaps(model), generation_policy_version: 1, generation_policies: ['greedy_v1'] }); +const caps = model => ({ ...legacyCaps(model), generation_policy_version: 1, generation_policies: ['greedy_v1'], + execution_error_version: 1 }); const result = (output, model) => ({ ...legacyResult(output, model), generation_policy: 'greedy_v1' }); const MODEL = 'qwen3-0.6b-v1'; @@ -101,7 +102,8 @@ test('SDK stream shape releases actual text and token usage only after terminal assert.equal(settled, false); assert.deepEqual(f.provider.observations, { submitted: 1, completed: 0, incomplete: 0, cleanup_confirmed: 0 }); assert.deepEqual(f.requests.map(row => row.operation.type), ['conversation_capabilities', 'submit_conversation']); - assert.deepEqual(f.requests[0].operation, { type: 'conversation_capabilities', generation_policy_version: 1 }); + assert.deepEqual(f.requests[0].operation, { type: 'conversation_capabilities', generation_policy_version: 1, + execution_error_version: 1 }); const input = f.requests[1].operation.conversation; assert.equal(input.generation_policy, 'greedy_v1'); assert.equal(input.visibility, 'private_local'); @@ -310,6 +312,39 @@ test('transient core failures keep their retryable status without claiming a com } }); +test('negotiated admitted budget exhaustion is cleanup-confirmed HTTP422, never a completed model result', async t => { + const f = await start(t, (socket, message) => { + reply(socket, message, 'admitted'); + reply(socket, message, 'error', {code: 'execution_budget_exceeded'}); + }, true); + try { + const response = await send(f.provider, request()); + assert.equal(response.status, 422); + assert.equal(JSON.parse(response.body).error.code, 'execution_budget_exceeded'); + assert.deepEqual(f.requests[0].operation, {type: 'conversation_capabilities', generation_policy_version: 1, + execution_error_version: 1}); + assert.equal(f.provider.diagnostics.summary.request_errors.execution_budget_exceeded, 1); + assert.deepEqual(f.provider.observations, {submitted: 1, completed: 0, incomplete: 0, cleanup_confirmed: 1}); + assert.deepEqual(f.provider.diagnostics.records, []); + } finally { await f.provider.close(); } +}); + +test('budget code before admission or under another correlation is not a cleanup receipt', async t => { + for (const mode of ['unadmitted', 'wrong-id']) { + const f = await start(t, (socket, message) => { + if (mode === 'wrong-id') reply(socket, message, 'admitted'); + reply(socket, mode === 'wrong-id' ? {...message, id: 'f'.repeat(32)} : message, + 'error', {code: 'execution_budget_exceeded'}); + }, true); + try { + const response = await send(f.provider, request()); + assert.notEqual(response.status, 422); + assert.equal(f.provider.observations.cleanup_confirmed, 0); + assert.equal(f.provider.diagnostics.summary.request_errors.execution_budget_exceeded, 0); + } finally { await f.provider.close(); } + } +}); + test('reaped execution failure followed by success accounts for both cleanups, not two model results', async t => { let attempts = 0; const f = await start(t, (socket, message) => { diff --git a/tests/opencode-config.test.cjs b/tests/opencode-config.test.cjs index b0099f2..e3bb04c 100644 --- a/tests/opencode-config.test.cjs +++ b/tests/opencode-config.test.cjs @@ -27,13 +27,30 @@ test('pinned runtime uses only core provider with separate one-shot tool boundar assert.equal(Object.hasOwn(env, 'HOME'), false); assert.equal(Object.hasOwn(env, 'OPENAI_API_KEY'), false); }); -test('legacy 0.6B defaults, prompts and permissions remain byte-identical', () => { +test('legacy 0.6B settings change only the explicit temperature capability', () => { const golden = JSON.parse(fs.readFileSync(path.join(__dirname, 'fixtures/opencode-default-settings.json'), 'utf8')); for (const cooperative of [false, true]) { + const expected = structuredClone(golden[cooperative ? 'cooperative' : 'default']); + assert.equal(Object.hasOwn(expected.config.provider.volparossa.models[MODEL], 'temperature'), false); + expected.config.provider.volparossa.models[MODEL].temperature = true; + expected.env.OPENCODE_CONFIG_CONTENT = JSON.stringify(expected.config); // Direct exact bytes, including key order and the synthetic environment; - // this regression fixture is not a password-storage or authentication hash. + // the historical fixture stays unchanged; no prompt, limit, permission or + // other environment change can hide behind this explicit capability delta. assert.equal(JSON.stringify(runtimeSettings({...input, cooperative})), - JSON.stringify(golden[cooperative ? 'cooperative' : 'default'])); + JSON.stringify(expected)); + } +}); +test('both model profiles advertise temperature so every coding role actually requests zero', () => { + for (const model of [MODEL, 'qwen3-4b-instruct-2507-v1']) { + for (const cooperative of [false, true]) { + const {config, env} = runtimeSettings({...input, model, cooperative}); + assert.equal(config.provider.volparossa.models[model].temperature, true); + for (const role of ['build', 'general', 'explore']) { + assert.equal(config.agent[role].temperature, 0); + } + assert.deepEqual(JSON.parse(env.OPENCODE_CONFIG_CONTENT), config); + } } }); test('4B settings use only the exact core-selected profile without changing task or tool policy', () => { diff --git a/tests/opencode-runtime.test.cjs b/tests/opencode-runtime.test.cjs index 6c63660..1607dd4 100644 --- a/tests/opencode-runtime.test.cjs +++ b/tests/opencode-runtime.test.cjs @@ -6,9 +6,26 @@ const assert = require('node:assert/strict'); const {spawn} = require('node:child_process'); const {PassThrough, Writable} = require('node:stream'); const {configuration, ownedOpenCode} = require('../src/opencode-runtime.cjs'); -const {readFrames, writeFrame, emptyProviderDiagnostic, emptyTaskDiagnostic} = require('../src/opencode-bridge.cjs'); +const {readFrames, writeFrame, emptyProviderDiagnostic, emptyTaskDiagnostic, + validProviderDiagnostic} = require('../src/opencode-bridge.cjs'); const READY = {type: 'ready', version: 1, execution: 'private_local', confidentialRemoteAvailable: false}; const RESULT = {text: 'Synthetic answer.', commands: 1, nativeTurnCompleted: true, taskVerified: false}; +test('provider counter schema accepts exact historical receipts and the new bounded terminal code', () => { + const current = emptyProviderDiagnostic(); + current.request_errors.execution_budget_exceeded = 1; + assert.equal(validProviderDiagnostic(current), true); + const old = structuredClone(current); + delete old.request_errors.execution_budget_exceeded; + assert.equal(validProviderDiagnostic(old), true); + assert.equal(Object.hasOwn(old.request_errors, 'execution_budget_exceeded'), false); + for (const changed of [true, -1, 65536, 'PRIVATE_CANARY']) { + const invalid = structuredClone(current); + invalid.request_errors.execution_budget_exceeded = changed; + assert.equal(validProviderDiagnostic(invalid), false); + } + old.request_errors.private_detail = 'PRIVATE_CANARY'; + assert.equal(validProviderDiagnostic(old), false); +}); function child(t, program) { const process = spawn(require('node:process').execPath, ['-e', program], {stdio: ['pipe', 'pipe', 'pipe'], env: {PATH: '/usr/bin:/bin', LANG: 'C.UTF-8', ELECTRON_RUN_AS_NODE: '1'}}); diff --git a/tests/opencode-session.test.cjs b/tests/opencode-session.test.cjs index c7bf932..f0fb074 100644 --- a/tests/opencode-session.test.cjs +++ b/tests/opencode-session.test.cjs @@ -19,7 +19,8 @@ function fixture(Task, receive, options = {}) { observations: options.observations ?? {submitted: 0, cleanup_confirmed: 0}, async close() { observed.push('provider-close'); if (options.badCleanup) throw Error('private detail'); }}; const hooks = {Task, preflight: async () => options.capabilities ?? - {...caps(options.model ?? DEFAULT_MODEL), generation_policy_version: 1, generation_policies: ['greedy_v1']}, + {...caps(options.model ?? DEFAULT_MODEL), generation_policy_version: 1, generation_policies: ['greedy_v1'], + execution_error_version: 1}, provider: async value => { binding.provider = value.model; return provider; }, prepare(cooperative) { assert.equal(cooperative, options.cooperative ?? false); }, cooperative: () => options.cooperative ?? false, port: async () => 1235, approvalMs: 15, @@ -85,10 +86,11 @@ test('owner-selected validated 4B core binds provider, native catalog and task b test('incompatible, widened or quarantined core fails before provider or native startup', async () => { const model = 'qwen3-4b-instruct-2507-v1'; for (const changed of [{quarantined: true}, {local_only: false}, {max_prompt_tokens: 262144}, - {generation_policies: []}, {model_profile: 'unreviewed-model'}]) { + {generation_policies: []}, {execution_error_version: undefined}, {execution_error_version: 2}, + {model_profile: 'unreviewed-model'}]) { class Task { constructor() { assert.fail('must not create task'); } } const f = fixture(Task, () => assert.fail('must not signal ready'), {capabilities: { - ...caps(model), generation_policy_version: 1, generation_policies: ['greedy_v1'], ...changed, + ...caps(model), generation_policy_version: 1, generation_policies: ['greedy_v1'], execution_error_version: 1, ...changed, }}); assert.equal(await f.done, 1); assert.deepEqual(f.binding, {}); assert.deepEqual(f.observed, []); @@ -124,7 +126,7 @@ test('owner cleanup accepts an actual provider retry after a correlated reaped f reply(socket, request, 'admitted'); if (++attempts === 1) reply(socket, request, 'error', {code: 'execution_failed'}); else reply(socket, request, 'result', {result: coreResult(undefined, model)}); - }, {...caps(model), generation_policy_version: 1, generation_policies: ['greedy_v1']}); + }, {...caps(model), generation_policy_version: 1, generation_policies: ['greedy_v1'], execution_error_version: 1}); const provider = await startChatCompletionsProvider({socketPath: core.socketPath, model, diagnostics: true}); class Task { async run() { diff --git a/tests/private-compute.test.cjs b/tests/private-compute.test.cjs index 85a7624..4a942b2 100644 --- a/tests/private-compute.test.cjs +++ b/tests/private-compute.test.cjs @@ -8,7 +8,7 @@ const net = require('node:net'); const os = require('node:os'); const path = require('node:path'); const { test } = require('node:test'); -const { PrivateCompute } = require('../src/private-compute.cjs'); +const { PrivateCompute, terminalCleanupConfirmed } = require('../src/private-compute.cjs'); function caps(overrides = {}) { return { visibility: 'private_local', local_only: true, model_profile: 'smollm2-360m-v1', @@ -78,6 +78,20 @@ async function fixture(t, handler, options = {}) { return { client, socketPath, directory, requests, sockets }; } +test('legacy Q&A never accepts the conversational terminal budget vocabulary', async t => { + const f = await fixture(t, (socket, request) => { + reply(socket, request, 'admitted'); + reply(socket, request, 'error', {code: 'execution_budget_exceeded'}); + }); + await f.client.connect(); + await assert.rejects(f.client.ask({question: 'Q', context: 'C'}), error => { + assert.equal(error.code, 'invalid_response'); + assert.equal(terminalCleanupConfirmed(error), false); + return true; + }); + assert.deepEqual(f.requests[0].operation, {type: 'capabilities'}); +}); + test('real Unix framing preserves the exact private result across partial and coalesced frames', async t => { const original = answer(); const f = await fixture(t, (socket, request) => { diff --git a/tests/private-conversation.test.cjs b/tests/private-conversation.test.cjs index 6241677..5a8b1fa 100644 --- a/tests/private-conversation.test.cjs +++ b/tests/private-conversation.test.cjs @@ -23,6 +23,80 @@ test('actual framed socket uses separate handshake and yields exact cleanup-conf await assert.rejects(f.client.ask({ question: 'legacy', context: 'must not send' })); }); +test('execution error negotiation is independent, strict, and unchanged when absent', async t => { + for (const value of [null, 0, 2, true]) { + assert.throws(() => new PrivateConversation('/fixture/core.sock', {executionErrorVersion: value}), + /unsupported_execution_error_version/); + } + for (const advertised of [caps(), {...caps(), execution_error_version: null}, + {...caps(), execution_error_version: 2}]) { + const f = await fixture(t, () => assert.fail('must not submit'), advertised, {executionErrorVersion: 1}); + await assert.rejects(f.client.connect(), {code: 'incompatible_capabilities'}); + assert.deepEqual(f.requests[0].operation, {type: 'conversation_capabilities', execution_error_version: 1}); + } + const unsolicited = await fixture(t, () => assert.fail('must not submit'), {...caps(), execution_error_version: 1}); + await assert.rejects(unsolicited.client.connect()); +}); + +test('only a negotiated admitted task has a budget-error cleanup receipt, with request-local provenance', async t => { + let submitted = 0; + const f = await fixture(t, (socket, request) => { + if (++submitted === 1) reply(socket, request, 'admitted'); + reply(socket, request, 'error', {code: 'execution_budget_exceeded'}); + }, {...caps(), execution_error_version: 1}, {executionErrorVersion: 1}); + await f.client.connect(); + let accepted; + await assert.rejects(f.client.submit(input()), error => { + accepted = error; + assert.equal(error.code, 'execution_budget_exceeded'); + assert.equal(terminalCleanupConfirmed(error), true); + return true; + }); + await assert.rejects(f.client.submit(input()), error => { + assert.equal(error.code, 'invalid_response'); + assert.equal(terminalCleanupConfirmed(error), false); + return true; + }); + assert.equal(terminalCleanupConfirmed({...accepted, cleanupConfirmed: true}), false); + assert.equal(terminalCleanupConfirmed(Object.assign(Error(accepted.message), {code: accepted.code})), false); + assert.equal(terminalCleanupConfirmed(accepted), true); +}); + +test('legacy, wrong-id, cancel-id and handshake budget codes never become cleanup receipts', async t => { + for (const mode of ['legacy', 'wrong-id', 'cancel-id']) { + let admitted; + const ready = new Promise(resolve => { admitted = resolve; }); + const f = await fixture(t, (socket, request) => { + if (request.operation.type === 'submit_conversation') { + reply(socket, request, 'admitted'); + if (mode === 'cancel-id') { admitted(); return; } + } + reply(socket, mode === 'wrong-id' ? {...request, id: 'f'.repeat(32)} : request, + 'error', {code: 'execution_budget_exceeded'}); + }, mode === 'legacy' ? caps() : {...caps(), execution_error_version: 1}, + mode === 'legacy' ? {} : {executionErrorVersion: 1}); + await f.client.connect(); + const controller = new AbortController(); + const pending = f.client.submit(input(), {signal: controller.signal}); + const rejected = assert.rejects(pending, error => { + assert.equal(error.code, 'invalid_response', mode); + assert.equal(terminalCleanupConfirmed(error), false, mode); + return true; + }); + if (mode === 'cancel-id') { await ready; controller.abort(); } + await rejected; + } + // Pure dispatch guard at the handshake boundary; no synthetic execution claim. + const client = new PrivateConversation('/fixture/core.sock', {executionErrorVersion: 1}); + client.state = 'connecting'; client.handshake = {id: 'a'.repeat(32)}; + assert.throws(() => client._response({version: 1, id: client.handshake.id, event: 'error', + code: 'execution_budget_exceeded'}), error => { + assert.equal(error.code, 'invalid_response'); + assert.equal(terminalCleanupConfirmed(error), false); + return true; + }); +}); + test('negotiated greedy requests require matching result evidence and cannot silently use a legacy core', async t => { const model = 'qwen3-0.6b-v1'; const negotiated = { ...caps(model), generation_policy_version: 1, generation_policies: ['greedy_v1'] }; diff --git a/tests/real-opencode-provider-errors.test.cjs b/tests/real-opencode-provider-errors.test.cjs new file mode 100644 index 0000000..0459950 --- /dev/null +++ b/tests/real-opencode-provider-errors.test.cjs @@ -0,0 +1,36 @@ +// SPDX-License-Identifier: GPL-3.0-only +// Synthetic frame bindings only; this test never starts OpenCode or a model. +'use strict'; +const assert = require('node:assert/strict'); +const {test} = require('node:test'); +const {syntheticResult} = require('./real_opencode_provider_errors.cjs'); +const {caps, input} = require('./conversation-fixture.cjs'); +const {validateResult} = require('../src/private-conversation.cjs'); +const MODEL = 'qwen3-0.6b-v1'; +const negotiated = {...caps(MODEL), generation_policy_version: 1, generation_policies: ['greedy_v1'], + execution_error_version: 1}; + +test('native error fixture does not invent policy from negotiated support', () => { + const conversation = input(); + for (const reason of ['invalid_output', 'wire_truncated']) { + const response = syntheticResult({type: 'incomplete', reason}, conversation); + assert.equal(Object.hasOwn(response, 'generation_policy'), false); + assert.equal(validateResult(response, negotiated, conversation), response); + assert.throws(() => validateResult({...response, generation_policy: 'greedy_v1'}, negotiated, conversation), + {code: 'invalid_conversation'}); + } +}); + +test('native error fixture retains the exact explicitly selected greedy policy', () => { + const conversation = {...input(), generation_policy: 'greedy_v1'}; + for (const reason of ['invalid_output', 'wire_truncated']) { + const response = syntheticResult({type: 'incomplete', reason}, conversation); + assert.equal(response.generation_policy, 'greedy_v1'); + assert.equal(validateResult(response, negotiated, conversation), response); + const {generation_policy: removed, ...missing} = response; + assert.equal(removed, 'greedy_v1'); + assert.throws(() => validateResult(missing, negotiated, conversation), {code: 'generation_policy_mismatch'}); + } + assert.throws(() => syntheticResult({type: 'assistant', text: 'synthetic'}, + {...conversation, generation_policy: 'unknown'})); +}); diff --git a/tests/real_opencode_provider_errors.cjs b/tests/real_opencode_provider_errors.cjs index 421eeee..25c1df7 100644 --- a/tests/real_opencode_provider_errors.cjs +++ b/tests/real_opencode_provider_errors.cjs @@ -12,6 +12,14 @@ const {caps, result, reply, fixture} = require('./conversation-fixture.cjs'); const MODEL = 'qwen3-0.6b-v1'; const hash = value => createHash('sha256').update(value).digest('hex'); +function syntheticResult(output, conversation) { + // Negotiating support does not select a policy. Mirror only the policy on + // the actual request, just as the core's checked result binding requires. + const selected = Object.hasOwn(conversation, 'generation_policy'); + if (selected) assert.equal(conversation.generation_policy, 'greedy_v1'); + return {...result(output, MODEL), ...(selected ? {generation_policy: conversation.generation_policy} : {})}; +} + async function terminalCase(reason, config) { const project = await fs.mkdtemp(path.join(os.tmpdir(), 'volparossa-opencode-error-')); await fs.chmod(project, 0o700); @@ -21,25 +29,37 @@ async function terminalCase(reason, config) { const abort = new AbortController(); const timer = setTimeout(() => abort.abort(), 90000); let runtime, fixtureError, taskError, submissions = 0, codingSubmissions = 0, approvals = 0; + const generationPolicies = {unspecified: 0, greedy_v1: 0}; let phase = 'fixture'; try { const core = await fixture({after: action => cleanups.push(action)}, (socket, message) => { try { assert.equal(message.operation.type, 'submit_conversation'); assert.equal(message.operation.conversation.visibility, 'private_local'); + const conversation = message.operation.conversation; + const selectedPolicy = Object.hasOwn(conversation, 'generation_policy'); + if (selectedPolicy) assert.equal(conversation.generation_policy, 'greedy_v1'); + generationPolicies[selectedPolicy ? 'greedy_v1' : 'unspecified']++; assert.ok(++submissions <= 8, 'bounded fixture requests'); const coding = message.operation.conversation.tools.some(tool => tool.name === 'bash'); - if (coding) assert.equal(++codingSubmissions, 1, 'terminal output must not cause native regeneration'); + if (coding) { + assert.equal(conversation.generation_policy, 'greedy_v1', 'native coding must select its configured policy'); + assert.equal(++codingSubmissions, 1, 'terminal output must not cause native regeneration'); + } // Titles and other upstream no-tool requests are not the coding task. - const output = coding ? {type: 'incomplete', reason} : {type: 'assistant', text: 'Synthetic retry check'}; reply(socket, message, 'admitted'); - reply(socket, message, 'result', {result: result(output, MODEL)}); + if (coding && reason === 'execution_budget_exceeded') { + reply(socket, message, 'error', {code: reason}); + } else { + const output = coding ? {type: 'incomplete', reason} : {type: 'assistant', text: 'Synthetic retry check'}; + reply(socket, message, 'result', {result: syntheticResult(output, conversation)}); + } } catch (error) { fixtureError ??= error; socket.destroy(); abort.abort(); } - }, caps(MODEL)); + }, {...caps(MODEL), generation_policy_version: 1, generation_policies: ['greedy_v1'], execution_error_version: 1}); phase = 'native_start'; runtime = await OpenCodeRuntime.start({...config, socketPath: core.socketPath}, {workspace: project}); phase = 'native_task'; @@ -59,18 +79,20 @@ async function terminalCase(reason, config) { assert.equal(codingSubmissions, 1); assert.equal(approvals, 0); const diagnostics = runtime.diagnostics; - assert.equal(diagnostics.incomplete_reasons[reason], 1); - assert.equal(diagnostics.request_errors.invalid_model_output, 1); + const budget = reason === 'execution_budget_exceeded'; + if (!budget) assert.equal(diagnostics.incomplete_reasons[reason], 1); + assert.equal(diagnostics.request_errors.invalid_model_output, budget ? 0 : 1); + assert.equal(diagnostics.request_errors.execution_budget_exceeded, budget ? 1 : 0); assert.equal(diagnostics.submitted, submissions); assert.equal(diagnostics.cleanup_confirmed, submissions); - assert.equal(diagnostics.incomplete, 1); + assert.equal(diagnostics.incomplete, budget ? 0 : 1); assert.equal(await fs.readFile(path.join(project, 'README.txt'), 'utf8'), marker); assert.deepEqual(await fs.readdir(project), ['README.txt']); phase = 'session_close'; await runtime.close(); runtime = null; return {reason, coding_submissions: codingSubmissions, auxiliary_submissions: submissions - codingSubmissions, - native_retry_count: 0, terminal_error: taskError.code, approvals, + native_retry_count: 0, terminal_error: taskError.code, approvals, generation_policies: generationPolicies, original_project_unchanged: true, session_cleanup_confirmed: true, provider_diagnostics: diagnostics}; } catch (error) { process.stderr.write(JSON.stringify({reason, phase, coding_submissions: codingSubmissions, @@ -113,7 +135,7 @@ async function main(args = process.argv.slice(2)) { const config = {version: 1, opencode: report.binary, opencodeSha256: report.binary_sha256, buildReport, node, nodeSha256: hash(await fs.readFile(node))}; const cases = []; - for (const reason of ['invalid_output', 'wire_truncated']) cases.push(await terminalCase(reason, config)); + for (const reason of ['invalid_output', 'wire_truncated', 'execution_budget_exceeded']) cases.push(await terminalCase(reason, config)); const evidence = {version: 1, native_runtime: true, runtime_version: report.runtime_version, source_commit: report.source_commit, binary_sha256: report.binary_sha256, node_sha256: config.nodeSha256, provider_sha256: hash(await fs.readFile(path.join(__dirname, '../src/chat-completions-provider.cjs'))), @@ -128,4 +150,4 @@ if (require.main === module) main().catch(error => { process.stderr.write(`Native retry check failed: ${String(error.message).slice(0, 200)}\n`); process.exitCode = 1; }); -module.exports = {main}; +module.exports = {main, syntheticResult}; diff --git a/tests/test_smoke_opencode_inference.py b/tests/test_smoke_opencode_inference.py index e31cacc..d8ef9f3 100644 --- a/tests/test_smoke_opencode_inference.py +++ b/tests/test_smoke_opencode_inference.py @@ -47,7 +47,7 @@ def test_exact_synthetic_inventory_is_bound_without_clean_git_claim(self): def test_profiles_are_closed_and_default_resources_are_unchanged(self): self.assertEqual(TRIAL.trial_profile(), dict(model_profile='qwen3-0.6b-v1', - core_revision='845cc84d0d0b766ab1c5227231dbf6c8eaeb8cc3', guest_memory_mib=6144, + core_revision='6a517b576baa17e7329661ee0476d1848081d114', guest_memory_mib=6144, core_memory_bytes=5 * TRIAL.GIB, qemu_memory_bytes=7 * TRIAL.GIB, host_available_bytes=8 * TRIAL.GIB, provision_budget_bytes=5 * TRIAL.GIB, scratch_gib=18, memory_failure='host_available_memory_below_8GiB')) @@ -57,7 +57,10 @@ def test_profiles_are_closed_and_default_resources_are_unchanged(self): def test_larger_profile_is_separate_and_requires_a_real_pin(self): self.assertRegex(TRIAL.LARGE_CORE, r'^[0-9a-f]{40}$') - self.assertNotEqual(TRIAL.LARGE_CORE, TRIAL.CORE) + # One reviewed core supports both profiles and the new negotiated error + # contract. Separate models/resources do not require artificial source forks. + self.assertEqual(TRIAL.LARGE_CORE, '6a517b576baa17e7329661ee0476d1848081d114') + self.assertNotEqual(TRIAL.LARGE_MODEL, TRIAL.MODEL) with patch.object(TRIAL, 'LARGE_CORE', None): with self.assertRaisesRegex(ValueError, 'larger_core_not_pinned'): TRIAL.trial_profile(TRIAL.LARGE_MODEL)