feat(guizang-product-video-skill): 默认确认动效与镜头推拉,新增逐句 Gemini 画外音(1.1.0) - #35
Merged
hwj123hwj merged 3 commits intoSep 21, 2026
Merged
Conversation
added 3 commits
September 21, 2026 16:00
…r-line voiceover - ask once for style, UI motion + camera zoom in/out, and whether to narrate - record the answers in plan.motion and voiceoverRequired with exception reasons - add scripts/make_voiceover.py: per-line EasyRouter Gemini TTS, 48 kHz assembly, evidence - mix the voice as a third layer, duck music under speech, -14 LUFS when narrated - validate narration windows and motion consistency in check_delivery.py, extend regressions - document the motion language and the measured EasyRouter TTS model list
…e voiceover gain real - missing motion/voiceoverRequired now warns instead of failing, so 1.0.1-era plans keep passing - mix the assembled audio.voiceover.file with its own gain; per-line files stay evidence and timing - apply lines[].gain while assembling and drop the per-line file requirement from the gate - align the narration wording in SKILL.md with the warning-level gate - add a real narrated-mix regression: gain reaches the stem, -14 LUFS, voice duck window
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
改了什么
guizang-product-video-skill升到 1.1.0:ask_user_question一次问清风格、动效(界面操作动效 + 镜头 zoom in/out)、要不要旁白;答案写进plan.json的motion.uiEffects、motion.cameraMove、voiceoverRequired,用户否决的一项必须留motion.exceptionReason/voiceoverExceptionReason。已确认或改片时不再重复问。references/story-and-copy.md):界面动效对应真实状态变化;镜头缩放幅度约 1.0→1.04–1.08、一镜一到两次;缩放后按变换后的实际位置查裁切;全静态不算错误。scripts/make_voiceover.py,走 EasyRouter(OpenAI 兼容)的 Gemini TTS 按镜头逐句生成,量出实际时长后拼成 48 kHz 单声道轨,并写evidence/voiceover.json(含可粘回 plan 的 planSnippet)。assets/voice-stem.wav。marketplace.json升到 1.1.0,packageKey 同步为..._1.1.0.zip。验证
ruby scripts/validate_skills.rb→Skill gate passed: 32 skill(s), 32 unique name(s).python3 -m unittest discover -s skills/guizang-product-video-skill/tests→ 18 tests, OK(新增 5 个用例:旁白理由、旁白缺文件/逐句证据、旁白超出镜头、动效理由、镜头与用户选择冲突)。/v1/models返回 102 个模型,可用 TTS 为gemini-3.1-flash-tts-preview与gpt-4o-mini-tts;不带-preview的gemini-3.1-flash-tts返回model_not_found(已写入文档)。/audio/speech+ Aoede 生成成功(一句 3.36 s,mean -20.1 dB / max -4.5 dB),端到端三轨混音输出正常,让位窗口与 -14 LUFS 目标符合预期,check_delivery --mix-report返回 ok。说明
EASYROUTER_API_KEY或视频工程.env),key 不写入 plan、证据或仓库;缺 key 时脚本明确报错,不做静默降级。SKILL.mdfrontmatter 的upstreamSha仍指向 op7418 上游的旧提交;本次是在镜像副本上的增强,是否重新对齐上游请维护者判断。