Source
Independent APPLAYLIST Skill Tester / adversarial QA pass over canonical Human DJ Review R1 + Human Preference Calibration R2.
Tester verdict
EXPERIMENT_VALIDITY=FAIL
The current real R4 calibration protocol mixes curation quality with uncontrolled DJ transition execution.
Finding 1 — transition metrics are not reproducible
The Bundle 63 reviewer workspace presents only Plan A / Plan B track ordering while requiring numeric scores for:
transition_smoothness
phrase_alignment
It does not bind a canonical transition execution recipe such as:
- outgoing/incoming mix windows
- bar length
- tempo handling
- EQ policy
- FX policy
- stems policy
- cue/loop behavior
- rendered reference transition
When the DJ performs each transition differently, those ratings measure a mixture of APPLAYLIST transition potential and reviewer execution skill/style. They are therefore not reproducible ground truth for APPLAYLIST transition quality.
Finding 2 — curation calibration label is contaminated
Bundle 68 requires the complete six-dimension HumanDJReview and resolves the single overall human preference directly against the Bundle 67 Competitive Curation shadow challenger.
Because the overall preference may be influenced by uncontrolled transition execution, the calibration can mark a curation challenger correct/incorrect for reasons outside the curation model being evaluated.
Required separation
CURATION_REVIEW != TRANSITION_FEASIBILITY_REVIEW != HUMAN_EXECUTION_REVIEW
A. Curation Review
Suitable for current A/B track-order cases:
- energy_flow
- dramaturgical_fit
- set_coherence
- alternative_usefulness
- track-selection / sequence preference
- confidence
B. Transition Feasibility Review
Must evaluate APPLAYLIST's proposed transition, not arbitrary DJ execution. Candidate dimensions:
- phrase-window quality
- energy handoff
- spectral compatibility
- vocal collision risk when explicit vocal evidence exists
- tempo/key feasibility where relevant
- transition-strategy suitability
This requires either a deterministic proposed mix recipe or a rendered/standardized preview.
C. Human Execution Review
Optional separate experiment. Only compare execution quality when mix instructions are standardized or the experiment explicitly studies the DJ's execution choices.
Protocol changes required before R4 calibration can be considered valid
- Split curation preference from transition preference.
- Do not require
transition_smoothness / phrase_alignment for curation-only calibration unless transition execution is standardized.
- Add
not_assessable / missing-evidence semantics per transition dimension instead of forcing a fabricated 1..5 score.
- Prevent Bundle 68 curation calibration from consuming a preference contaminated by transition-execution dimensions.
- Preserve tie/abstain.
- Preserve qualitative notes as evidence rather than coercing them into numeric ratings.
- Add tests proving curation calibration cannot depend on human execution-only ratings.
- Keep optimizer activation / production / PDM training authority false.
Current R4 implication
The existing qualitative track-selection review remains useful curation evidence.
The current 12-case full 6D review should be PAUSED as calibration evidence until this protocol boundary is corrected. No historical qualitative feedback should be fabricated into missing numeric fields.
Authority
IMPLEMENTATION_AUTHORIZATION=NO
MERGE_AUTHORIZATION=NO
OPTIMIZER_RANKING_ACTIVATION=NO
RELEASE_AUTHORIZATION=NO
DEPLOY_AUTHORIZATION=NO
PRODUCTION_ACTIVATION=NO
PDM_TRAINING=NO
Source
Independent APPLAYLIST Skill Tester / adversarial QA pass over canonical Human DJ Review R1 + Human Preference Calibration R2.
Tester verdict
EXPERIMENT_VALIDITY=FAILThe current real R4 calibration protocol mixes curation quality with uncontrolled DJ transition execution.
Finding 1 — transition metrics are not reproducible
The Bundle 63 reviewer workspace presents only Plan A / Plan B track ordering while requiring numeric scores for:
transition_smoothnessphrase_alignmentIt does not bind a canonical transition execution recipe such as:
When the DJ performs each transition differently, those ratings measure a mixture of APPLAYLIST transition potential and reviewer execution skill/style. They are therefore not reproducible ground truth for APPLAYLIST transition quality.
Finding 2 — curation calibration label is contaminated
Bundle 68 requires the complete six-dimension HumanDJReview and resolves the single overall human
preferencedirectly against the Bundle 67 Competitive Curation shadow challenger.Because the overall preference may be influenced by uncontrolled transition execution, the calibration can mark a curation challenger correct/incorrect for reasons outside the curation model being evaluated.
Required separation
CURATION_REVIEW != TRANSITION_FEASIBILITY_REVIEW != HUMAN_EXECUTION_REVIEWA. Curation Review
Suitable for current A/B track-order cases:
B. Transition Feasibility Review
Must evaluate APPLAYLIST's proposed transition, not arbitrary DJ execution. Candidate dimensions:
This requires either a deterministic proposed mix recipe or a rendered/standardized preview.
C. Human Execution Review
Optional separate experiment. Only compare execution quality when mix instructions are standardized or the experiment explicitly studies the DJ's execution choices.
Protocol changes required before R4 calibration can be considered valid
transition_smoothness/phrase_alignmentfor curation-only calibration unless transition execution is standardized.not_assessable/ missing-evidence semantics per transition dimension instead of forcing a fabricated 1..5 score.Current R4 implication
The existing qualitative track-selection review remains useful curation evidence.
The current 12-case full 6D review should be PAUSED as calibration evidence until this protocol boundary is corrected. No historical qualitative feedback should be fabricated into missing numeric fields.
Authority
IMPLEMENTATION_AUTHORIZATION=NO
MERGE_AUTHORIZATION=NO
OPTIMIZER_RANKING_ACTIVATION=NO
RELEASE_AUTHORIZATION=NO
DEPLOY_AUTHORIZATION=NO
PRODUCTION_ACTIVATION=NO
PDM_TRAINING=NO