Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,25 @@ the [releases page](https://github.com/JeremySNR/cutawan/releases).
This project uses [semantic versioning](https://semver.org/), loosely: while
still pre-1.0, minor bumps carry new features and patch bumps carry fixes.

## [0.13.0] - 2026-09-25

### Added

- Optional editorial ranking (beta) reviews retained speech, surrounding context and sampled frames using common hook, clarity, value, payoff and audience-fit criteria. It marks uncertainty, defers repeated ideas, preserves the old ranking for offline comparison, and shows reasons and evidence in the editor. Extra analysis calls are shown before enabling it.
- Optional **Find visual moments (beta)** scans sampled frames across the source before selecting clips, then inspects promising demonstrations, reveals and reactions more closely. It can suggest clips without spoken audio. Completed scans are cached by source content, transcript, provider/model and instructions; failed scans remain retryable.
- Visual discovery reports show successful and failed source sections, sampling gaps and rejected proposals. The scan uses at most twelve additional analysis calls plus provider retries. Sparse frames can miss brief events; this is not continuous-video understanding or a measured improvement in engagement.
- Offline quality comparisons now support blind publishability decisions, measured repair time, severe defects, reviewer coverage and source-separated summaries. Missing clips and missing reviews remain visible, with a protocol for collecting untouched human-labelled holdouts.

### Improved

- Score badges distinguish editorial assessments from legacy scores. Changed source selections require review again; a high score is not presented as a prediction of virality or certification of the export.
- Visual candidates preserve their observed action and crossing speech boundaries, start with conservative framing, and keep pause removal and auto zoom off. Candidates that cannot fit completely are rejected. Final overlap removal runs after visual-payoff repair.

### Validation

- Automated regressions use scripted model responses; they do not establish human preference or superiority over OpusClip. Run the documented human comparison workflow before making those claims.
- A bounded live ChatGPT check exercised editorial review, visual discovery and refinement. The API route was not live-tested. See `docs/provider-smoke.md` for the recorded results and limits.

## [0.12.1] - 2026-09-23

### Fixed
Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,8 +77,9 @@ On the default API route, a typical estimate is **~$0.36/hour of video** for Whi

- **Import anything.** Local files (MP4/MOV/MKV/WEBM and more) or paste a URL from YouTube, Vimeo, TikTok, Twitch, or any site yt-dlp supports. Private or SSO-protected videos (like enterprise Vimeo) work by borrowing the login from your browser. No server integration needed.
- **Whisper transcription** with word-level timestamps. Long videos are chunked automatically and checkpointed, so retries and re-generations never pay for transcription twice.
- **Find visual moments (beta).** Opt in to sampled source-wide discovery before transcript selection, including demonstrations and visible events without speech. The app shows sampling gaps and failed coverage. [How it works and its limits](docs/source-discovery.md).
- **Viral moment detection backed by research.** An LLM picks self-contained hook, build, payoff micro-stories (not clips that trail off mid-setup). You can steer it with your own prompt if you want, like "find the funniest exchanges". A second AI pass reviews every clip ending and extends it to the beat that actually completes the thought.
- **Two-pass virality scoring (0-99).** A text rubric based on Berger and Milkman's *What Makes Online Content Viral?* (JMR 2012), plus measured vocal energy, combined with a vision pass from Kayal et al. (ACL 2025) that scores sampled frames for scroll-stopping potential.
- **Editorial selection and ranking.** Review retained speech, nearby context and sampled frames against common hook, clarity, value, payoff and audience-fit criteria. Incomplete stories need evidence; uncertainty stays visible, repeated ideas move down the list, and scores are provisional editorial assessments. [How ranking and its comparison baseline work](docs/editorial-ranking.md).

**Making them good**

Expand Down
17 changes: 17 additions & 0 deletions benchmarks/editorial-ranking/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Fixed-pool editorial ranking experiment

Status: **the legacy recipe is frozen; no human quality result is recorded here.**

`legacy-recipe.json` was captured from the working tree before adding editorial reranking. It records the exact legacy score blend and final overlap-selection functions, source-file hashes, the Git base, and the presence of uncommitted changes. Do not replace this recipe when tuning a challenger. Introduce a new experiment version instead.

The question is whether a new ordering of the **same candidate pool** puts more complete, useful, distinct clips near the top and reduces correction work. This is narrower than whether a complete new pipeline beats a historical release or OpusClip. Discovery misses, candidates filtered by earlier repair/review, and different model proposal responses are outside this comparison.

Freeze the candidate pool before editorial selection. Preserve complete clip snapshots, including IDs, original insertion order, provisional scores, source boundaries, cut/protection ranges, framing state and captions. Preserve the transcript and source identity too: saving a clip's start/end alone cannot reproduce its rendered result after caption or framing edits.

Use the recorded legacy rule to obtain the baseline ordering. For the challenger, use the editorial decisions made from that exact pool. Both arms must use the same source bytes, frozen transcript, rendering inputs and top-five budget. Later user edits or automatic framing must not silently replace saved benchmark inputs. Exports remain unavailable until actual videos have been rendered from those inputs.

Before comparing outputs, declare the source groups, audience instructions, top-K budget, clip-length constraint, code/model/configuration revisions, primary metric and failure policy. Rank-only selection does not establish discovery recall or a reliable probability of virality. Do not tune from held-out review labels.

At least two independent editors should review anonymized exports and then compare their source context. Use the [quality benchmark](../../docs/quality-benchmark.md) for publishable/repairable/reject decisions, severe defects, measured repair seconds, missing-output coverage and source-group summaries. Do not fabricate labels to turn a saved snapshot into a completed experiment.

The primary comparison is publishable clips per requested top-five slot, accompanied by source-level results, source-fidelity defects, diversity and actual repair time. Report fallback/unreviewed cases and empty slots. Keep per-reviewed-clip publishability alongside full-slot yield so abstaining from difficult cases cannot conceal lost coverage.
37 changes: 37 additions & 0 deletions benchmarks/editorial-ranking/legacy-recipe.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
{
"schemaVersion": 1,
"id": "legacy-editorial-v1",
"frozenAt": "2026-09-25T18:07:34.970Z",
"gitHead": "15c0694d4195b4376b7a163aa83affbaff668291",
"workingTree": "Includes uncommitted source-discovery increment; gitHead alone does not identify these source files.",
"purpose": "Fixed algorithmic baseline for within-candidate-pool editorial ranking. This is a protocol, not a completed experiment or human judgment.",
"stage": "After existing visual review, attempted payoff repair, and known incomplete/incoherent candidate removal; before final score ordering and overlap removal.",
"input": "A deep snapshot of every surviving Clip, including original provisional blended viralityScore, source ranges and complete render-relevant edit state. Preserve original candidate insertion order for ties.",
"score": "When visual review succeeds: clamp(round(0.6 * proposalScore + 0.4 * visualScore), 0, 99). When unavailable: retain proposalScore. These are uncalibrated model judgments, not engagement probabilities.",
"selection": "Stable descending viralityScore order. Iterate in that order; reject a candidate only when its suggested interval overlaps any retained interval by more than 0.4 of the shorter suggested duration. Do not enforce a top-K quota or pad empty slots.",
"controlledComparison": "Both arms use the identical surviving input pool, original transcript, source bytes and frozen rendering inputs. Any retrieval/prompt change is shared by both arms and therefore not evaluated by this ranking-only experiment.",
"limitations": [
"The historical proposal score is not modality calibrated.",
"The pool already excludes candidates rejected by upstream review, so this cannot measure discovery recall or full old-versus-new pipeline quality.",
"Source-file hashes identify the inspected working tree but do not imply that current model responses are reproducible.",
"No measured quality change, human ratings, OpusClip outputs, or engagement outcomes are recorded by this artifact."
],
"sourceEvidence": {
"src/main/pipeline/highlights.ts": {
"sha256": "e12ebf0cf9a4b63791ab35a930b4f14d8c17a023d4ebb1d64cb2456714206c55"
},
"src/main/pipeline/visualScore.ts": {
"sha256": "df2dc90af03dd4bbbd55c61f9795b1f465f3fac259a0e3455cdb7be1f4270834"
},
"src/main/pipeline/visualCandidates.ts": {
"sha256": "c82d8420483b54f3d3296d3f4d4c03d74f05ef764ae13784152393a1bf4e4fa8"
},
"src/main/pipeline/index.ts": {
"sha256": "5af2bd4eb85c27eeb9e68f173d8fd6de05e9ca38d2605e8797e3a55868452d7a"
}
},
"frozenFunctions": {
"ensembleScore": "export function ensembleScore(textScore: number, visualScore: number): number {\n return Math.max(0, Math.min(99, Math.round(TEXT_WEIGHT * textScore + (1 - TEXT_WEIGHT) * visualScore)))\n}",
"dedupeClips": "export function dedupeClips(clips: Clip[], maxOverlapFraction = 0.4): Clip[] {\n const kept: Clip[] = []\n for (const clip of clips) {\n const dur = clip.suggestedEnd - clip.suggestedStart\n const tooSimilar = kept.some((k) => {\n const overlap =\n Math.min(clip.suggestedEnd, k.suggestedEnd) - Math.max(clip.suggestedStart, k.suggestedStart)\n if (overlap <= 0) return false\n const kDur = k.suggestedEnd - k.suggestedStart\n return overlap / Math.min(dur, kDur) > maxOverlapFraction\n })\n if (!tooSimilar) kept.push(clip)\n }\n return kept\n}"
}
}
53 changes: 53 additions & 0 deletions benchmarks/holdout/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# Untouched quality holdout protocol

Status: **protocol and offline tooling implemented; footage, human labels, and matched OpusClip outputs are not yet collected for this holdout.** Nothing in this directory is a measured competitive result. The nine public videos under `benchmarks/public-corpus` have already guided development and remain regression cases, regardless of their older provisional split labels.

## Freeze the questions before acquiring footage

Use three separate experiments. Do not change the question after seeing which system wins.

1. **Discovery:** give every system the identical full source, destination, duration constraints and five-clip budget. Preserve its original suggestion ordering. Compare top-five publishability and distinct worthwhile reference moments retrieved. Include no-good-moment sources; they reveal false-positive selection and should be reported as their own stratum. Low yield can be appropriate on those sources.
2. **Editing:** give every system identical human-selected source intervals and matched visual/caption settings. Compare source fidelity, composition, timing, sound and measured correction work. A system unable to accept fixed intervals is unavailable for this experiment, not a zero-error pass.
3. **Complete workflow:** start from identical full sources using each product's documented recommended settings. Record import-to-first-usable-clip time, processing failures, retries, repair time, cost and publishable output. Keep product defaults and manual interventions in provenance.

Record product version/model, settings, prompts, date, hardware, provider, source SHA-256, export SHA-256 and all manual changes before review. Freeze the code revision and protocol before generating test exports. Never pick a system's best rerun without counting other attempts. Predeclare how transient retries are handled.

## Acquire genuinely new source groups

Recruit or license recordings that have not been used to tune Cutawan, including footage from creators outside the existing corpus. Obtain permission for local testing and, separately, any provider upload involved in generating competitive exports. Keep media out of Git and retain acquisition/permission records next to its inventory. This protocol does not authorize uploading private footage.

Start with a feasibility pilot using development footage, then aim for at least 30 untouched source recordings from at least 15 independent creator groups. These are planning targets, not a power calculation or a guarantee of statistical confidence; use pilot variance to determine the sample needed for the decision. Reserve sufficient footage for later release confirmation rather than consuming the entire holdout during iteration.

Include expert interviews and narrated demonstrations as primary strata. Include remote calls, multiple/overlapping speakers, varied accents and speaking styles, soft speech, background noise/music, moving subjects, scene changes, off-screen questions, small text, moving insets, long setup/payoff chains and weak/no-good-moment sources. Predeclare any languages or genres outside the claim. Do not present English interviews as proof of sports/gameplay or multilingual quality.

Assign stable opaque `creatorGroup` and `recordingGroup` IDs before splitting. The same creator, host/channel family, studio/layout or recording session stays on one side of development/test; union overlapping relationships conservatively. Re-encodes and excerpts of a recording inherit its group. The runner rejects cross-split reuse of supplied groups and identical source hashes. It warns when groups are missing for compatibility, so incomplete metadata is not a certified holdout.

Maintain a private inventory with case ID, source path/hash/duration, creator and recording groups, acquisition date, permissions, tags, language, capture setup, and a record of whether developers have inspected it. Limit access to the untouched test labels. Once a source is used to tune a fix, treat it as regression footage and evaluate the fix on new held-out groups.

## Annotate the entire source independently

Two editors watch the full recording **before seeing any system's proposals**. Each identifies worthwhile self-contained moments, their necessary context, promised payoff, source interval and rationale. Include valuable visual events with little or no useful transcript: a visible result, demonstration, reaction or readable comparison. Annotate hard negatives such as intros, advertisements, repeated ideas and incomplete promises. Explicitly record sources with no worthwhile moment; do not force three or five positives.

Reconcile independent source annotations into distinct moment references without inspecting system outputs. Preserve both original annotations and the adjudication record. A missing reference must remain a possible annotation gap: interval-IoU recall measures agreement with those references, not exhaustive creative merit. Record reasons for disagreements and unresolved alternatives. When evaluating timing or speakers, annotate all relevant speech and overlap, not just easy excerpts.

Write validated `QualitySample` reference JSON and a version 1 benchmark manifest using the [runner contract](../../docs/quality-benchmark.md). References must be frozen before running discovery. Use source hashes and optional source paths to verify identical media. For each product, store outputs under their original rank with explicit `rank`; missing exports remain missing, and a run that returns nothing still includes that case with `highlights: []`. Never remove failed or difficult cases.

## Review exports and time corrections

Run `npm run quality:benchmark -- compare <manifest> <new-directory>`. A coordinator retains the private maps and metrics and sends only the blind HTML and anonymized videos to reviewers. Inspect the package for branding leaks before the experiment. Randomize independently for repeat rounds; avoid presenting paired alternatives consecutively where recognition would bias first impressions.

At least two reviewers independently rate each clip using the generated form. Watch the export standalone, then check the original source for missing context and changed meaning. Mark publishable/repairable/reject, severe defects, source verification, dimension scores and timestamped notes. Time corrections on a copy of accepted exports in a fixed editing environment; include verification time and distinguish editing from machine waiting in separate run notes. Blank timing means unmeasured. Reviewers should not discuss their decisions until submissions are frozen.

Aggregate with `summarize-review`. Keep conditional publishability alongside full top-K yield, reviewed/eligible coverage, rendering failures, missing outputs, severe defects and repair-timing coverage. At least two files are necessary but do not prove reviewer independence. Resolve disagreements with a third reviewer or a documented adjudication, retaining original votes; the current tool conservatively requires unanimous publishable votes and exposes disagreements rather than silently adjudicating them.

## Proposed release gates

Freeze gates after the development pilot and before the untouched comparison. Initial targets for the primary interview/demo strata are:

- At least four publishable suggestions per five requested, while also reporting no-good-moment strata and abstention behavior separately.
- Median measured repair time below 60 seconds for accepted clips, with high, reported timing coverage and tail times inspected.
- At least 90% of reviewed clips actually labelled Ready are publishable as shown; report all declared-Ready coverage so selective omissions cannot satisfy the gate.
- No observed changed-meaning or seriously damaged-speech defect in a release candidate; adjudicate every such flag and report sample size. Zero observed does not prove zero risk.
- No loss on previously fixed regression examples, and no material failure rate increase on declared primary strata.

Use confidence intervals and source-level cases alongside aggregates. The runner currently provides source-cluster intervals for observed publishable yield, not a powered superiority test or direct pairwise-preference test. Add a preregistered matched preference study with independent source-level uncertainty before claiming that Cutawan beats OpusClip. Reviewer preference and publishability do not establish audience lift; that requires later authorized publishing experiments with comparable audiences and exposure.
Loading
Loading