Skip to content

Improve Jev synthetic data auditing and selection provenance - #21

Merged
sileod merged 8 commits into
mainfrom
improve-jev-synthetic-audit
Sep 29, 2026
Merged

sileod merged 8 commits into
mainfrom
improve-jev-synthetic-audit

Conversation

@sileod

@sileod sileod commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Why

The synthetic Jev pipeline already preserves raw generation/Jev responses, but two research-validity gaps remain:

  1. scenario quality is checked before Jev labeling, while confident teacher mistakes are not independently surfaced;
  2. when downsampling with --n-target, confidence-bucket selection can make the retained distribution conditional on Jev's own uncertainty.

What this PR changes

  • adds an optional independent post-Jev teacher audit:

    • auditor sees only the generated state/questions, never Jev probabilities;
    • it independently answers Choice / Score / Noul questions with a confidence;
    • local code compares its answer against the Jev distribution;
    • high-confidence Jev/auditor disagreements are explicitly flagged;
    • raw auditor responses are content-addressed and cached;
    • filtering disagreements is opt-in, while the default keeps flagged rows auditable.
  • adds a teacher-independent selection reservoir:

    • a configurable fraction of a downsampled corpus is chosen from state IDs only;
    • the rest follows the requested Jev ambiguity buckets;
    • every selected bundle records selection_reason and selection_bucket.
  • exports provenance on flat training rows:

    • target origin / returned teacher model;
    • Jev max probability and entropy;
    • selection arm/bucket;
    • independent-auditor agreement/confidence/disagreement flag.
  • adds an example audited config:

    • DeepSeek generation via Albert;
    • pinned Jev 1.13 targets via OpenRouter;
    • independent OpenAI/Luna teacher audit.

Intended research use

This does not replace Jev targets with another LLM's judgments. The independent model is a diagnostic layer. The goal is to make confident teacher failures measurable and to enable matched ablations between unfiltered and confidence-curated synthetic data.

Tests

Adds offline tests for:

  • confident teacher-disagreement detection;
  • audit caching;
  • teacher-independence of the reservoir;
  • audited-config provider separation.

CI is expected to run the full suite on Python 3.10 and 3.13.

@sileod
sileod marked this pull request as ready for review September 29, 2026 12:01
@sileod
sileod merged commit d759df7 into main Sep 29, 2026
2 checks passed
@sileod
sileod deleted the improve-jev-synthetic-audit branch September 29, 2026 12:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant