Skip to content

fix(guizang-product-video-skill): 按旁白答案条件上报 EASYROUTER_API_KEY - #36

Merged
hwj123hwj merged 1 commit into
OrionStarAI:mainfrom
hwj123hwj:fix/guizang-video-voiceover-prereq
Sep 21, 2026
Merged

hwj123hwj merged 1 commit into
OrionStarAI:mainfrom
hwj123hwj:fix/guizang-video-voiceover-prereq

Conversation

@hwj123hwj

Copy link
Copy Markdown
Collaborator

问题

环境检查脚本从不提 EASYROUTER_API_KEY,而唯一叫 agent 去索取它的地方写在 references/onboarding.md 里——按流程,这份文档只在检查报出缺项时才读。

于是形成一个自锁:永远报不出来的缺项,永远读不到它的说明。实际后果是一次完整走完的旁白制作:画面全部做完、混音前才发现没有 key,用户白等一整轮。

按 --voiceover 的答案条件化上报,而不是一律要求 key——多数片子不要旁白,那时不该多出任何提示。

改动

scripts/check_environment.py

  • 新增 --voiceover yes|no|undecided,默认 no。
  • 只有 yes 时把 key 报成 prereqs(不是 missing):凭据是用户给的,不是环境坏了,也不该让引擎检查(浏览器启动)被跳过。
  • undecided 只给一条告警。
  • 不读 plan.json:第 2 步时 plan 还是起步工程的技术样片,voiceoverRequired 是过期值,读它会得到错误答案、并静默压掉提示。这个决定只存在于对话里,所以必须用参数传。
  • next 现在把两类缺项都说出来,不互相遮蔽。
  • 报告的 voiceover 字段只含 requested / keyResolved / keySource,不含 key 明文,也进了指纹,所以 key 状态变化会触发重新上报。

SKILL.md:选「要旁白」的那一刻就解析并索取 key(它挡混音不挡渲染);检查时把答案传进去;写稿前先量本片真实语速。

references/audio-sourcing.md:两处修正,都来自实测。

  1. 「中文约每秒 6–9 字是预警线」是阅读约束,不是旁白目标。实测旁白约 4.2 字/秒(十句 3.5–4.8),照阅读线排稿会系统性超时(实测第一版 10 句里 7 句越界)。
  2. 明确「不要靠加速」的三种手段都实测过、都不要用:提示词语速指令(4.64 → 6.97 字/秒)拿到的是压缩式加速,听感会被识破;speed 在 Gemini 后端被静默忽略、在 gpt-4o-mini-tts 上精确生效;atempo 只是兜底。另外记下语速不能跨句外推(同一音色在探测句 5.56、在本片正式稿 4.21 字/秒)。

references/onboarding.md:两个实测过的环境问题——macOS 系统 python3(3.9 + LibreSSL)访问网关 TLS 失败、/audio/speech 偶发首次请求被对端断开(应重试,不是 key 或配额问题)。

tests/test_regressions.py:两个新回归——四种决策组合下的上报行为,以及 key 值不进入报告或 evidence/environment.json。

验证

  • ruby scripts/validate_skills.rb → Skill gate passed: 32 skill(s), 32 unique name(s)
  • python3 -m unittest discover -s tests → 21 tests OK(原 19)
  • 手工核对四种组合:no 静默 / yes 无 key → prereqs + ready:false / undecided → 仅告警 / yes + .env → 从 .env 解析
  • 用假 key 验证 evidence/environment.json 里只有 keySource,无明文

版本 1.1.0 → 1.1.1(marketplace.json 的 version 与 packageKey 同步)。

… actually needed

The environment check never mentioned EASYROUTER_API_KEY, and the only place that
told the agent to ask for it sat in references/onboarding.md — which the flow gates
behind "load it only for reported gaps". A gap that is never reported is never read
about, so a narrated film kept discovering the missing key at the mixing step, after
the entire picture had been built.

- check_environment.py: add `--voiceover yes|no|undecided`. Only `yes` (the step-1
  answer) reports the key, in a separate `prereqs` list so a missing credential is
  not read as a broken toolchain and does not skip the browser launch; `undecided`
  warns instead. The answer is deliberately not read from plan.json — at that point
  the plan is still the starter sample and its voiceoverRequired is stale, which
  would silently suppress the reminder.
- SKILL.md: resolve the key the moment the user says yes (it blocks mixing, not
  rendering); pass the answer to the check.
- audio-sourcing.md: give the narration pace positively — use the voice's natural
  pace, measured on this film's own copy (typically 4–6 chars/s), and adjust by
  changing the voice or shortening the line. The existing 6–9 chars/s figure is a
  *reading* line, not a narration target, and budgeting copy to it overruns the shots.
- onboarding.md: Homebrew python3 for the gateway (system 3.9 + LibreSSL fails TLS)
  and the transient first-request disconnect.
- tests: two regressions for the gating, plus one asserting the key value never
  reaches the report or evidence/environment.json.
@hwj123hwj
hwj123hwj merged commit a41af71 into OrionStarAI:main Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant