perf(agentx): refresh Kimi-K3 MI355X LMCache curve / 刷新 Kimi-K3 MI355X LMCache 曲线 - #2804
perf(agentx): refresh Kimi-K3 MI355X LMCache curve / 刷新 Kimi-K3 MI355X LMCache 曲线#2804hyukjlee wants to merge 6 commits into
Conversation
Use the merged Kimi-K3 recipe with vLLM nightly 7c5dc571 and LMCache 0.5.5rc3 for C1, C8, C14, and TP8/DCP8 C40. Keep the validated C40 profile with GPU memory utilization reduced to 0.88. Assisted-by: OpenAI Codex <codex@openai.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
2 similar comments
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Assisted-by: OpenAI Codex <codex@openai.com>
Assisted-by: OpenAI Codex <codex@openai.com>
为 Kimi-K3 MI355X LMCache 曲线添加 C44 和 C48,并沿用 C40 的 TP8/DCP8、无推测解码和 GMU 0.88 配置。 Assisted-by: OpenAI Codex <codex@openai.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=33614613326 |
将 LMCache MP 心跳超时提高到 90 秒,并将 worker 回收超时提高到 300 秒,避免 C14 大批量传输期间的短暂阻塞触发错误恢复路径。 Assisted-by: OpenAI Codex <codex@openai.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=33618719560 |
将 C1、C8 和 C14 的 gpu-memory-utilization 从 0.90 降至 0.88,使全部六个测试点使用相同的 GMU。移除上一轮未证实的 heartbeat workaround,以便单独验证显存利用率变化。 Assisted-by: OpenAI Codex <codex@openai.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=33630491494 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=33631260867 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=33631260867 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=33631260867 |
Description / 说明
Refreshes the canonical
kimik3-fp4-mi355x-vllm-agentic-mtpcurve while preserving the tuning established by merged PR perf(agentx): retune Kimi-K3 FP4 MI355X vLLM agentic MTP recipe #2787.Bumps vLLM ROCm to
nightly-7c5dc571cbd1064ecc8a9b1045637ff647aa22cb.Pins LMCache to
0.5.5rc3+rocm7.2and enables LMCache DRAM KV offload for C1, C8, C14, C40, C44, and C48.Keeps C1/C8/C14 on TP8 with DSpark MTP and sets
--gpu-memory-utilization 0.88for all six points.Uses the C40 TP8/DCP8 no-spec profile for C40/C44/C48: LMCache chunk size 12288, DCP KV-cache interleave size 1536, max-num-seqs 80, 16K batched tokens, and full CUDA graphs through size 4096.
刷新标准的
kimik3-fp4-mi355x-vllm-agentic-mtp曲线,并保留已合并 PR perf(agentx): retune Kimi-K3 FP4 MI355X vLLM agentic MTP recipe #2787 中验证过的调优配置。将 vLLM ROCm 镜像更新到
nightly-7c5dc571cbd1064ecc8a9b1045637ff647aa22cb。固定 LMCache 版本为
0.5.5rc3+rocm7.2,并为 C1、C8、C14、C40、C44 和 C48 启用 LMCache DRAM KV offload。C1/C8/C14 保持 TP8 和 DSpark MTP 配置,全部六个测试点统一使用
--gpu-memory-utilization 0.88。C40/C44/C48 使用相同的 TP8/DCP8 无投机解码配置:LMCache chunk size 12288、DCP KV-cache interleave size 1536、max-num-seqs 80、16K batched tokens,以及最大到 4096 的完整 CUDA graph。
Duplicate-work check / 重复工作检查
This does not duplicate open PR #2795, which removes LMCache and targets a different no-offload C52 profile. It also supersedes the old LMCache 0.5.4rc work in PR #2598 by updating the already-merged canonical PR #2787 recipe to LMCache rc3 and the new vLLM image.
本 PR 不与开放的 PR #2795 重复;该 PR 移除了 LMCache,并针对不同的无 offload C52 配置。本 PR 也取代了 PR #2598 中旧的 LMCache 0.5.4rc 工作,将已合并 PR #2787 的标准配置更新到 LMCache rc3 和新的 vLLM 镜像。
Type of Change / 变更类型
Checklist / 检查清单
perf-changelog.yamlwithout editing historical entries / 每项影响性能的基准变更均已在perf-changelog.yaml文件末尾追加记录,且未修改历史条目/reuse-sweep-runafter a final all-green sweep with evals passing / 通过复用方式合并前,授权维护者需在最终完整 sweep 和 eval 全部通过后评论/reuse-sweep-runAI assistance was used to prepare and validate this change. The submitting human must review every changed line and confirm the benchmark results before merge.
本变更使用了 AI 辅助。提交者在合并前必须审阅每一处改动,并确认基准测试结果。
Note
Low Risk
Benchmark and sweep configuration only; no production runtime or auth paths change, though refreshed numbers will replace the published Kimi-K3 MI355X curve.
Overview
Refreshes the
kimik3-fp4-mi355x-vllm-agentic-mtpAgentX curve: bumps the vLLM ROCm image tonightly-7c5dc571, pins LMCache to0.5.5rc3+rocm7.2, and records the sweep inperf-changelog.yaml.amd-master.yamlnarrows the MTP + LMCache DRAM offload matrix to C1, C8, C14 and adds TP8/DCP8 points at C40, C44, C48 withspec-decoding: none. A thinkimik3_fp4_mi355x.shwrapper setsSPEC_DECODING=noneand delegates to the main script.kimik3_fp4_mi355x_mtp.shkeys tuning on${SPEC_DECODING}:${CONC}: MTP arms usegpu-memory-utilization0.88; high-concurrency no-spec runs get a dedicated server profile (fixedmax-num-seqs80, expanded CUDA-graph capture through 4096,--stream-interval,--prefix-match-unit128, prefill FP8 quant in attention config, extra compilation custom ops). LMCache uses chunk-size 12288, exports backend metadata, and a lighter import sanity check; with DCP and LMCache,--cp-kv-cache-interleave-size1536,ROCM_AITER_MLA, and related env flags replace the prior TRITON-only DCP path.Reviewed by Cursor Bugbot for commit a720a04. Bugbot is set up for automated code reviews on this repo. Configure here.