Skip to content

feat(guizang-product-video-skill): 默认确认动效与镜头推拉,新增逐句 Gemini 画外音(1.1.0) - #35

Merged
hwj123hwj merged 3 commits into
OrionStarAI:mainfrom
hwj123hwj:feat/guizang-video-motion-voiceover
Sep 21, 2026
Merged

hwj123hwj merged 3 commits into
OrionStarAI:mainfrom
hwj123hwj:feat/guizang-video-motion-voiceover

Conversation

@hwj123hwj

Copy link
Copy Markdown
Collaborator

改了什么

guizang-product-video-skill 升到 1.1.0:

  • 需求确认从一问变三问:新片默认用 ask_user_question 一次问清风格、动效(界面操作动效 + 镜头 zoom in/out)、要不要旁白;答案写进 plan.json 的 motion.uiEffects、motion.cameraMove、voiceoverRequired,用户否决的一项必须留 motion.exceptionReason / voiceoverExceptionReason。已确认或改片时不再重复问。
  • 动效语言(references/story-and-copy.md):界面动效对应真实状态变化;镜头缩放幅度约 1.0→1.04–1.08、一镜一到两次;缩放后按变换后的实际位置查裁切;全静态不算错误。
  • 逐句画外音:新增 scripts/make_voiceover.py,走 EasyRouter(OpenAI 兼容)的 Gemini TTS 按镜头逐句生成,量出实际时长后拼成 48 kHz 单声道轨,并写 evidence/voiceover.json(含可粘回 plan 的 planSnippet)。
  • 混音:人声为锚,音乐在语音期间让位(默认 -6 dB / attack 0.15 s / release 0.4 s,hold 用该句实测时长);有旁白时响度目标 -14 LUFS,无旁白仍 -16;新增输出 assets/voice-stem.wav。
  • 交付检查:有旁白需文件 + 逐句证据 + 混音证据;无旁白需理由;旁白超出所属镜头或比镜头长会告警;镜头声明与用户选择冲突会报错。
  • 文档与模板同步:README、audio-sourcing、audio-and-qa、onboarding、init_project、agents/openai.yaml;marketplace.json 升到 1.1.0,packageKey 同步为 ..._1.1.0.zip。

验证

  • ruby scripts/validate_skills.rb → Skill gate passed: 32 skill(s), 32 unique name(s).
  • python3 -m unittest discover -s skills/guizang-product-video-skill/tests → 18 tests, OK(新增 5 个用例:旁白理由、旁白缺文件/逐句证据、旁白超出镜头、动效理由、镜头与用户选择冲突)。
  • 真实链路实测(2026-09-21):网关 /v1/models 返回 102 个模型,可用 TTS 为 gemini-3.1-flash-tts-preview 与 gpt-4o-mini-tts;不带 -preview 的 gemini-3.1-flash-tts 返回 model_not_found(已写入文档)。/audio/speech + Aoede 生成成功(一句 3.36 s,mean -20.1 dB / max -4.5 dB),端到端三轨混音输出正常,让位窗口与 -14 LUFS 目标符合预期,check_delivery --mix-report 返回 ok。
  • 不带旁白的旧路径回归通过(-16 LUFS、无 voice-stem、无 voiceover 条目)。

说明

  • 旁白需要使用者自己的 EasyRouter key(EASYROUTER_API_KEY 或视频工程 .env),key 不写入 plan、证据或仓库;缺 key 时脚本明确报错,不做静默降级。
  • SKILL.md frontmatter 的 upstreamSha 仍指向 op7418 上游的旧提交;本次是在镜像副本上的增强,是否重新对齐上游请维护者判断。

weijian added 3 commits September 21, 2026 16:00
…r-line voiceover

- ask once for style, UI motion + camera zoom in/out, and whether to narrate
- record the answers in plan.motion and voiceoverRequired with exception reasons
- add scripts/make_voiceover.py: per-line EasyRouter Gemini TTS, 48 kHz assembly, evidence
- mix the voice as a third layer, duck music under speech, -14 LUFS when narrated
- validate narration windows and motion consistency in check_delivery.py, extend regressions
- document the motion language and the measured EasyRouter TTS model list
…e voiceover gain real

- missing motion/voiceoverRequired now warns instead of failing, so 1.0.1-era plans keep passing
- mix the assembled audio.voiceover.file with its own gain; per-line files stay evidence and timing
- apply lines[].gain while assembling and drop the per-line file requirement from the gate
- align the narration wording in SKILL.md with the warning-level gate
- add a real narrated-mix regression: gain reaches the stem, -14 LUFS, voice duck window
@hwj123hwj
hwj123hwj merged commit f14ed7f into OrionStarAI:main Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant