You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Run the Roadmap 06 Slice 2 live-model evaluation needed to decide whether Kiln should promote any cross-harness operator communication default, interaction profile, response skill combination, or prompt fallback.
Issue #77 delivered the provider-neutral contract and transport evidence. It intentionally promoted no global default and did not establish writing-quality improvement. This issue owns that missing behavioral evidence.
Roadmap 06 Slice 2 owns prompt-component evaluation and removal ablation.
docs/research/active/provider-neutral-communication.md owns current provider/harness evidence and this evaluation question.
docs/evaluations/communication-governance-v1.md verifies deterministic transport and attribution, not live writing quality.
clear-writing and action-first-communication are optional procedures, not universal defaults.
Operator instruction profiles are durable preferences; they do not prove equivalent model behavior.
Decision questions
Does a concise detail default reduce unnecessary output without losing required facts, evidence, warnings, verification, or residual risk?
Does an explicit operator communication profile reduce repetitive or formulaic assistant patterns across Codex, Claude Code, and OpenCode?
Do clear-writing or action-first-communication add measurable value after the profile is present, or only duplicate prompt content?
Which native controls, instruction projections, or prompt fallbacks work by exact model and harness revision?
Do parent and child responses preserve the same required-content obligations when their communication intent resolves independently?
Which candidates regress exact formats, tool use, findings-first reviews, long-form explanations, or multilingual responses?
Candidate conditions
Evaluate at least:
provider/harness default;
operator communication instruction profile only;
profile plus native detail control where supported;
profile plus action-first-communication;
profile plus clear-writing;
profile plus both skills only if the separate candidates justify the combined cell;
component-removal ablations for every candidate considered for promotion.
Claude output style and OpenCode per-turn provider options remain unsupported unless fresh capability evidence proves ownership, exact transport, and rollback. Do not silently approximate them.
Fixture set
Include simple factual answers, repository findings, implementation closeout, blocked work, critical review, exact JSON, multi-question prompts, Spanish and English turns, requested long-form teaching, safety/authority warnings, managed-child handoffs, and malformed-tool recovery.
Use fresh sessions and freeze model, harness/adapter revision, task prompt, repository/config identity, skill revision, profile revision, reasoning policy, tool catalog, and output budget.
Measurements
Primary:
required-fact and required-content recall;
unsupported claims and false completion;
exact-format compliance;
correctness and deterministic task outcome where applicable;
human-rated comprehension and actionability with a blinded rubric.
Secondary:
output tokens and time to first useful information;
repeated claims and restatement rate;
generic praise, ceremonial framing, redundant summaries, excessive headings/lists, unnecessary next-action sections, and routine process narration;
latency, cost, tool-trajectory changes, and route-specific regressions.
The stylistic-pattern rubric must be revisioned and independently scored on a calibration set. Phrase counting alone is not a quality metric. Concision cannot compensate for correctness, evidence, safety, or tool-use regression.
Promotion gate
Define thresholds before collecting the promotion run. A candidate may be promoted only for exact model/harness revisions where it improves the primary outcome or is non-inferior on correctness while materially improving operator comprehension or information density. Aggregate gains cannot hide a material route-specific regression.
Unsupported routes remain unsupported or explicitly omitted. A local operator preference may remain active without being promoted as a Kiln product default.
Acceptance criteria
Baseline, candidates, and removal ablations are replayable by prompt manifest, config, profile, skill, model, route, and harness identity.
Codex, Claude Code, and OpenCode receive fresh-session trials for every admitted mechanism.
Parent and managed-child behavior is evaluated separately.
The rubric distinguishes verbosity from formulaic writing and from required evidence.
Human scoring reports agreement and unresolved ambiguity.
Failed, blocked, omitted, and capability-unsupported trials remain visible.
A promotion decision is scoped by model/harness revision and records semantic loss.
Any promoted default updates communication research, evaluation record, architecture, global-config guidance, descriptors, projections, and rollback evidence together.
If no candidate clears the gate, no default is promoted and the negative result is retained.
Independent benchmark-readiness and communication review have no unresolved high or medium findings.
Non-goals
No universal personality prompt.
No blanket ban list treated as deterministic proof of human-quality prose.
No automatic skill admission on every turn without measured incremental value.
No overwrite of operator-owned Claude output styles or unmanaged harness settings.
No claim that matching response length implies matching quality.
No configuration-schema redesign; Roadmap 12 may expose the eventual admitted intent but does not own this evaluation.
Objective
Run the Roadmap 06 Slice 2 live-model evaluation needed to decide whether Kiln should promote any cross-harness operator communication default, interaction profile, response skill combination, or prompt fallback.
Issue #77 delivered the provider-neutral contract and transport evidence. It intentionally promoted no global default and did not establish writing-quality improvement. This issue owns that missing behavioral evidence.
Related work
docs/research/active/provider-neutral-communication.mdowns current provider/harness evidence and this evaluation question.docs/evaluations/communication-governance-v1.mdverifies deterministic transport and attribution, not live writing quality.clear-writingandaction-first-communicationare optional procedures, not universal defaults.Decision questions
clear-writingoraction-first-communicationadd measurable value after the profile is present, or only duplicate prompt content?Candidate conditions
Evaluate at least:
action-first-communication;clear-writing;Claude output style and OpenCode per-turn provider options remain unsupported unless fresh capability evidence proves ownership, exact transport, and rollback. Do not silently approximate them.
Fixture set
Include simple factual answers, repository findings, implementation closeout, blocked work, critical review, exact JSON, multi-question prompts, Spanish and English turns, requested long-form teaching, safety/authority warnings, managed-child handoffs, and malformed-tool recovery.
Use fresh sessions and freeze model, harness/adapter revision, task prompt, repository/config identity, skill revision, profile revision, reasoning policy, tool catalog, and output budget.
Measurements
Primary:
Secondary:
The stylistic-pattern rubric must be revisioned and independently scored on a calibration set. Phrase counting alone is not a quality metric. Concision cannot compensate for correctness, evidence, safety, or tool-use regression.
Promotion gate
Define thresholds before collecting the promotion run. A candidate may be promoted only for exact model/harness revisions where it improves the primary outcome or is non-inferior on correctness while materially improving operator comprehension or information density. Aggregate gains cannot hide a material route-specific regression.
Unsupported routes remain
unsupportedor explicitly omitted. A local operator preference may remain active without being promoted as a Kiln product default.Acceptance criteria
Non-goals