Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 65 additions & 0 deletions docs/jev/synthetic-pilot-2026-09-29.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,3 +38,68 @@ preflight. A 32-state generation comparison took 69 seconds with two keys
(16 requests per slot) and 89 seconds with one key. Neither short run showed
a rate-limit error, so this suggests a speed benefit but does not establish
independent sustained quotas.

## Broad interpretive text understanding

The audited and 4,000-state configurations add low-weight general reading skills
through the ordinary domain sampler. The current 14 skills cover topic, emotion,
communicative intent, claim support, stance, document purpose, main point,
implicit concern, intended audience, argument role, stakeholder perspective,
social implication, evidence strength, and message tone. They share generation,
validation, critic, Jev annotation, and independent audit with other skills.
Arithmetic, chronology, entity lookup, and literal retrieval are left to the
procedural segment. There are no benchmark labels or benchmark-specific paths.
Prompt files have stable names (`generate.txt`, `critic.txt`, and
`teacher_audit.txt`) rather than version suffixes.

A deterministic 4,000-spec draw yielded 5,890 questions, including 606 general
reading questions (10.3%). All 36 domains appeared. Each of the 14 general
skills appeared in at least 19 domains. The overall mix still includes existing
skills such as toxicity, sentiment, and groundedness. Difficulty levels 1–5
remain available; the general segment intentionally contains both simple text
classification and questions requiring several cues.

The first natural live pilot, `.synthetic_runs/natural_general_pilot_115/`, used
an earlier eight-skill candidate mix. It yielded 96 retained states and 123
Jev decisions from 115 generated states. All 18 retained general questions had
Jev/auditor agreement, but manual review found an entity-type question that
classified an issue rather than an entity. That, plus overlap with the
procedural generators, prompted the current interpretive mix.

The current natural live pilot is
`.synthetic_runs/interpretive_general_pilot_120/` (ignored by Git). Its 120
Albert generations produced 117 validated states, 94 critic-approved and
retained states, and 119 Jev decisions. Sixteen retained questions used nine
of the general reading skills. The independent auditor answered 118 of 119
questions, disagreed with Jev on 12 overall and one general question, flagged
34 questions for review, and found no confident disagreements; 93 of 94 states
passed audit. Manual review of the 16 general questions found useful variety,
but also drift: two `claim_support` questions asked for a main concern, an
`audience_inference` question became ticket routing, an `implicit_concern`
question guessed a customer's feelings without the customer's own words, and a
`practical_implication` question duplicated policy application. The last skill
has been replaced by `social_implication`, and the generation and critic prompts
now explicitly reject those drifts. Jev/auditor agreement did not catch them.

That current pilot took 582 seconds end to end with one Albert key and pilot
critic/auditor limits of 120 requests per minute: about 13,900 retained states
or 17,700 decisions per day if that rate holds. The earlier 115-state pilot
took 481 seconds with two Albert keys: about 17,200 retained states or 22,100
decisions per day by the same extrapolation. These are short-run estimates,
not sustained throughput guarantees; the production configs use 40 critic and
auditor requests per minute, and quota, retries, cost, and longer-run quality
may change throughput. Auditor agreement is diagnostic, not ground truth.

A focused quality probe, `.synthetic_runs/interpretive_skill_quality_probe_16/`,
sampled four specs each for claim support, intended audience, implicit concern,
and social implication from the ordinary 4,000-spec draw. Fourteen of 16 states
passed the original validator and critic, yielding 21 decisions. Manual review
found an ambiguous intended audience, a negotiation implication with several
plausible readings, and a question that projected today's idle crew into a
client visit tomorrow. The auditor flagged the first two but agreed on the
third. The critic now requires a per-question, exact state quote and explicit
skill, support, and unique-answer checks. On the same 15 validated states this
stricter DeepSeek check rejected the ambiguous negotiation question, although
it still missed the unwarranted projection about tomorrow. This is a concrete
remaining quality risk. A larger manual sample is needed before treating an
unattended 4,000-state run as clean training data.
10 changes: 7 additions & 3 deletions src/tasksource/jev/synthetic/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ class GenerationConfig:
concurrency: int = 20
max_output_tokens: int = 4000
seed: int = 42
prompt_version: str = "generate_v1"
prompt_version: str = "generate"
n_states: int = 1000


Expand All @@ -40,12 +40,16 @@ class SamplerConfig:
difficulty_weights: dict | None = None
ambiguity_weights: dict | None = None
distractor_weights: dict | None = None
# Optional low-weight general language skills enter the ordinary domain mix.
general_text_skill_weight: float = 0.0

def __post_init__(self):
from .schemas import FORMATS
unknown = set(self.question_formats) - set(FORMATS)
if unknown or not self.question_formats:
raise ValueError(f"question_formats keys must be among {FORMATS}: {sorted(self.question_formats)}")
if not math.isfinite(self.general_text_skill_weight) or self.general_text_skill_weight < 0:
raise ValueError("general_text_skill_weight must be finite and non-negative")
for name in ("difficulty_weights", "ambiguity_weights", "distractor_weights"):
weights = getattr(self, name)
if weights is not None and (not weights or any(float(v) < 0 for v in weights.values())
Expand All @@ -71,7 +75,7 @@ class CriticConfig:
# When `provider` is absent, it inherits the generator provider.
provider: ProviderConfig | None = None
model: str = "deepseek-v4-flash-0731"
prompt_version: str = "critic_v1"
prompt_version: str = "critic"
temperature: float = 0.0
requests_per_minute: int = 40

Expand Down Expand Up @@ -101,7 +105,7 @@ class TeacherAuditConfig:
# check, explicitly choose a different provider/model from the generator.
provider: ProviderConfig | None = None
model: str = "deepseek-v4-flash-0731"
prompt_version: str = "teacher_audit_v1"
prompt_version: str = "teacher_audit"
temperature: float = 0.0
requests_per_minute: int = 40
min_teacher_confidence: float = 0.85
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ generation:
concurrency: 10
max_output_tokens: 4000
seed: 42
prompt_version: generate_v2
prompt_version: generate
n_states: 1000

sampler:
Expand All @@ -34,7 +34,7 @@ critic:
base_url: https://albert.api.etalab.gouv.fr/v1
model: deepseek-v4-flash-0731
model: deepseek-v4-flash-0731
prompt_version: critic_v1
prompt_version: critic
temperature: 0.0

annotator:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ generation:
concurrency: 10
max_output_tokens: 4000
seed: 42
prompt_version: generate_v2
prompt_version: generate
n_states: 1000

sampler:
Expand All @@ -32,7 +32,7 @@ critic:
base_url: https://albert.api.etalab.gouv.fr/v1
model: deepseek-v4-flash-0731
model: deepseek-v4-flash-0731
prompt_version: critic_v1
prompt_version: critic
temperature: 0.0
requests_per_minute: 40

Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# DeepSeek generation + independent OpenAI-model teacher audit + pinned Jev targets.
# DeepSeek generation and critic + independent teacher audit + pinned Jev targets.
# This is intentionally more expensive than the overnight baseline. The auditor
# never sees Jev's probabilities; it independently answers each question and only
# flags high-confidence disagreements.
Expand All @@ -17,7 +17,7 @@ generation:
concurrency: 20
max_output_tokens: 4000
seed: 42
prompt_version: generate_v2
prompt_version: generate
n_states: 4000

sampler:
Expand All @@ -31,6 +31,7 @@ sampler:
difficulty_weights: {'1': 0.30, '2': 0.25, '3': 0.20, '4': 0.15, '5': 0.10}
ambiguity_weights: {minimal: 0.30, low: 0.25, moderate: 0.20, high: 0.15, extreme: 0.10}
distractor_weights: {'0': 0.35, '1': 0.30, '2': 0.20, '3': 0.15}
general_text_skill_weight: 0.3

critic:
enabled: true
Expand All @@ -41,7 +42,7 @@ critic:
base_url: https://albert.api.etalab.gouv.fr/v1
model: deepseek-v4-flash-0731
model: deepseek-v4-flash-0731
prompt_version: critic_v1
prompt_version: critic
temperature: 0.0
requests_per_minute: 40

Expand All @@ -61,7 +62,7 @@ teacher_audit:
base_url: https://openrouter.ai/api/v1
model: openai/gpt-4.1-mini
model: openai/gpt-4.1-mini
prompt_version: teacher_audit_v1
prompt_version: teacher_audit
temperature: 0.0
requests_per_minute: 40
min_teacher_confidence: 0.85
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ generation:
concurrency: 20
max_output_tokens: 4000
seed: 42
prompt_version: generate_v2
prompt_version: generate
n_states: 4000
sampler:
seed: 42
Expand All @@ -31,6 +31,7 @@ sampler:
difficulty_weights: {'1': 0.30, '2': 0.25, '3': 0.20, '4': 0.15, '5': 0.10}
ambiguity_weights: {minimal: 0.30, low: 0.25, moderate: 0.20, high: 0.15, extreme: 0.10}
distractor_weights: {'0': 0.35, '1': 0.30, '2': 0.20, '3': 0.15}
general_text_skill_weight: 0.3
critic:
enabled: true
provider:
Expand All @@ -40,7 +41,7 @@ critic:
base_url: https://albert.api.etalab.gouv.fr/v1
model: deepseek-v4-flash-0731
model: deepseek-v4-flash-0731
prompt_version: critic_v1
prompt_version: critic
temperature: 0.0
requests_per_minute: 40
annotator:
Expand Down
4 changes: 2 additions & 2 deletions src/tasksource/jev/synthetic/configs/mock_pilot.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ generation:
concurrency: 8
max_output_tokens: 4000
seed: 42
prompt_version: generate_v2
prompt_version: generate
n_states: 20

sampler:
Expand All @@ -32,7 +32,7 @@ critic:
base_url: mock://
model: mock-generator
model: mock-generator
prompt_version: critic_v1
prompt_version: critic
temperature: 0.0

annotator:
Expand Down
4 changes: 2 additions & 2 deletions src/tasksource/jev/synthetic/configs/openai_luna.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ generation:
concurrency: 20
max_output_tokens: 4000
seed: 42
prompt_version: generate_v1
prompt_version: generate
n_states: 1000

sampler:
Expand All @@ -32,7 +32,7 @@ critic:
base_url: https://api.openai.com/v1
model: gpt-6-luna
model: gpt-6-luna
prompt_version: critic_v1
prompt_version: critic
temperature: 0.0

annotator:
Expand Down
35 changes: 29 additions & 6 deletions src/tasksource/jev/synthetic/critic.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@
import asyncio
import hashlib
import json
import re
import time
from pathlib import Path

Expand Down Expand Up @@ -46,6 +47,31 @@ def mock_critique(bundle: dict) -> dict:
return {"pass": not issues, "issues": issues, "score": 1.0 if not issues else 0.0}


def checked_verdict(response: dict, bundle: dict) -> dict:
"""Require one grounded skill/answer check for every live critique."""
issues = [str(issue) for issue in response.get("issues", [])]
checks = response.get("checks")
questions = bundle.get("questions", [])
expected = {q["question_id"] for q in questions}
if not isinstance(checks, list) or len(checks) != len(questions) or {
check.get("question_id") for check in checks if isinstance(check, dict)
} != expected:
issues.append("missing or duplicate per-question critic checks")
else:
state = re.sub(r"\s+", " ", bundle.get("state", "")).casefold()
for check in checks:
qid = check["question_id"]
quote = check.get("evidence_quote")
if not isinstance(quote, str) or not quote.strip() or re.sub(
r"\s+", " ", quote).strip().casefold() not in state:
issues.append(f"{qid}: evidence quote not found in state")
for field in ("skill_match", "supported", "unique_answer"):
if check.get(field) is not True:
issues.append(f"{qid}: critic {field} check failed")
return {"pass": response.get("pass") is True and not issues,
"issues": issues, "score": float(response.get("score", 0.0))}


class RequestPacer:
"""Space critic calls so a batch stays below the provider's minute limit."""

Expand Down Expand Up @@ -86,13 +112,10 @@ async def _critique_one(sem, client, model: str, temperature: float,
await pacer.wait()
result = await providers.chat_complete(
client, model, [{"role": "user", "content": prompt}],
temperature=temperature, max_tokens=1000)
temperature=temperature, max_tokens=2000)
try:
verdict = extract_json_object(result["text"])
verdict = {"pass": bool(verdict.get("pass", False)),
"issues": list(verdict.get("issues", [])),
"score": float(verdict.get("score", 0.0))}
except (ValueError, json.JSONDecodeError, TypeError):
verdict = checked_verdict(extract_json_object(result["text"]), bundle)
except (ValueError, json.JSONDecodeError, TypeError, KeyError):
verdict = {"pass": False, "issues": ["unparseable critic response"], "score": 0.0}
cached.write_text(json.dumps(
{"cache_key": key, "state_id": bundle["state_id"],
Expand Down
51 changes: 51 additions & 0 deletions src/tasksource/jev/synthetic/prompts/critic.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
You are a strict quality critic for Jev state bundles (one state, one or more questions).
Valid formats are `choice` (one best option), `noul` (one yes/no proposition),
and `score` (an ordered scale). A bundle with one question is valid.

BUNDLE:
{{BUNDLE_JSON}}

Check:
1. All questions genuinely refer to the SAME state (no orphan/hallucinated references).
2. Questions are not paraphrases of each other; each tests a distinct aspect/skill.
3. Questions are not mutually inconsistent and do not leak the answer.
4. Requested distinct formats/skills are actually represented.
In particular, `claim_support` must assess evidence for a specific claim;
`audience_inference` must identify an intended audience, not ticket routing;
`implicit_concern` must be grounded in the concerned person's own words;
and `social_implication` must concern human response, not policy compliance.
5. Each yes/no question asks about a single proposition, and each score question
matches the meaning of its ordered rubric. Reject mismatches such as urgency
with likelihood levels or a numeric scale without defined endpoints.
6. The state belongs to its stated domain and does not mention Jev or the
dataset construction process.
7. Reject a question when the state omits the policy, deadline, current time,
or decision standard needed to answer it. Reject specialist safety or legal
decisions that require outside rules not supplied in the state.
8. For every choice question, derive the answer independently from the state,
then check every option against it. Reject if none matches, if two or more
match, or if the best option depends on an unstated interpretation. Do the
arithmetic for numeric options rather than relying on the wording.
9. Check temporal and numeric questions against the actual dates and numbers.
Reject an event-ordering question that treats a planned deadline as an event
that occurred, or asks for the order of events whose times are not stated.
Reject a comparison that assumes an unreported start, end, or total.
10. Reject reference questions with multiple plausible antecedents and relation
questions that infer a role or relationship absent from the state, even if
it seems plausible from tone or context.

For EACH question, record a short EXACT quote copied from the state that
supports its answer. Check that the question tests its declared skill and that
one answer is uniquely supported. A quoted sentence about today's conditions
does not establish tomorrow's conditions. A quote that merely makes one option
plausible is insufficient when another option is also plausible.
The skill list is extensible: judge whether each question matches its declared
skill, not whether the skill name appears in this prompt. Keep every issue
short and specific; do not put a reasoning transcript inside an issue string.

Output STRICT JSON only:
{"pass": bool, "issues": [str], "score": float,
"checks": [{"question_id": str, "evidence_quote": str,
"skill_match": bool, "supported": bool, "unique_answer": bool}]}
Include exactly one check for every question. Set pass to false if any check
fails. The evidence quote must be a contiguous substring of the state.
22 changes: 0 additions & 22 deletions src/tasksource/jev/synthetic/prompts/critic_v1.txt

This file was deleted.

Loading
Loading