Broaden synthetic Jev with interpretive reading tasks - #23
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
v2.Live check
A natural 120-state Albert run produced 117 validated states, 94 critic-approved states, and 119 Jev decisions in 582 seconds with one Albert key. The independent auditor answered 118/119 questions, disagreed on 12 overall and one of the 16 retained general questions, and reported no confident disagreements. The short-run rate extrapolates to about 17,700 retained decisions/day; production rate limits and sustained quota may change that.
Manual reading found skill drift and unsupported inferences that Jev/auditor agreement did not catch. A focused 16-state probe drove tighter skill instructions and the critic evidence checks. The stricter critic rejected an ambiguous negotiation implication, but missed a prediction about tomorrow based only on today's idle crew. The pilot note documents the remaining quality risk. This PR broadens the candidate distribution and improves screening; it does not certify an unattended 4,000-state export as clean training data.
Verification
PYTHONPATH=.:src pytest -q: 217 passedgit diff --check: clean.synthetic_runs/