Fix Kimi K3 B200 AIPerf metrics configuration - #2569
Conversation
中文:添加 B200 Kimi K3 AgentX 延迟、均衡吞吐、GPU 常驻高并发和 CPU KV 卸载配置,并通过 srt-slurm 启动多节点聚合式推理。
中文:绕过 B200 上不受支持的自定义集合通信,避免 TP16 在权重加载前停滞。
中文:修正 B200 DEP 启动参数、KV 事件发布、GPU 计数与 CPU DRAM 预算,并禁用跨节点不受支持的 FlashInfer 集合通信。
中文:为跨节点 TP16 和 TEP16 强制使用 PyNCCL,跳过会在 B200 集群上停滞的 MNNVL 自定义集合通信初始化。
中文:为跨节点 TP16 和 TEP16 禁用 Kimi K3 LatentMoE tail fusion,避免 torch symmetric memory 无法跨节点传递文件描述符。
中文:为跨节点 TP16 和 TEP16 禁用 allreduce/RMS fusion,避免 profile_run 再次选择不受支持的 FlashInfer MNNVL 工作区。
中文:改用可运行的 B200 TP8 x PP2 配置,并移除无法启动的跨节点 TP16、TEP16 和 DEP16 配置。
中文:DSpark 草稿模型不支持流水线并行,因此改回无 PP 的 TP16 配置,并保留 GPU 与 CPU KV 容量档位。
中文:将 main 合并到 B200 分支
中文:使用新版 Kimi nightly 镜像测试 B200
中文:禁用 Kimi 跨节点融合 latent-MoE 尾部路径,改用可移植的 PyNCCL 通信。
中文:强制 Kimi 跨节点 latent-MoE 使用可移植的 PyNCCL 通信路径。
中文:在两个 B200 配方中安装 Kimi 兼容性启动脚本。
中文:跳过跨节点 B200 不支持的 Kimi FlashInfer 对称内存工作区探测。
中文:强制 Kimi 跨节点通信使用 PyNCCL 可移植路径。
中文:补齐 Kimi 基准测试变更日志条目的末尾换行。
中文:移除 TP16 方案不再使用的 DEP 计数逻辑与测试,保持 PR 范围聚焦。
中文:将 DRAM 预留比例调整为 0.63,使每节点 1,889 GB 的资源预留覆盖每个 TP rank 220 GiB 的 KV 卸载池。
中文:移除已不适用于当前 TP16 方案的 DEP 启动器注释。
中文:并发 64 预检因 TTFT 指标覆盖率仅为 97.6% 而失败。将 DRAM 卸载容量档上限调整为待验证的并发 48。
中文:并发 64 预检已被工作流接受并上传完整聚合结果,保留该 DRAM 卸载容量端点以刻画饱和区间。
中文:合并 main 并解决冲突
中文:合并 main 后保留性能变更日志末尾换行
中文:合并最新 main,按追加规则保留 B200 Kimi K3 基准测试变更日志,并复用已通过的完整扫描结果。
中文:B200 改用上游 TP16+EP16 Kimi K3 推理路径,移除本地 vLLM 补丁,并为 GPU 常驻和 DRAM 卸载配置加入真实 block 验证评估。
中文:B200 双节点 TEP 配置改用上游 eager 模式,避免 CUDA graph capture 探测不可用的跨节点 MNNVL workspace。
中文:B200 双节点拓扑改为节点内 TP8、跨节点 DP2/EP16,使用上游 vLLM 原生路径并避免跨节点 FlashInfer MNNVL workspace。
中文:合并 main 并将 B200 Kimi K3 基准测试条目重新追加到 perf-changelog.yaml 末尾。
中文:缩短多节点评估产物名称,避免超过 GitHub Actions 的 256 字符上限。
为四个 Kimi K3 配方启用显式 AIPerf 服务端指标端点、提示词缓存明细,并要求导出包含 vllm: 前缀。
将 SimpleCPUOffloadConnector 搜索范围扩展至 c8 至 c64,以便与驻留曲线逐点比较并确定交叉点。
同步主分支并保留双方追加的性能变更记录。
移除固定 srt-slurm 不支持的字段,并通过自定义基准环境显式传递聚合 vLLM 指标地址。
同步已合并的 Kimi K3 提交,使后续修复仅包含受支持的指标环境变量变更。 # Conflicts: # benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8dp2-latency-dspark-agentic.yaml # benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8dp2-latency-dspark-eval-agentic.yaml # benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8dp2-vllm-simple-offload-dspark-agentic.yaml # benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8dp2-vllm-simple-offload-dspark-eval-agentic.yaml
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
追加记录固定 srt-slurm 的指标配置修复。
| command: bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh | ||
| env: | ||
| INFMAX_CONTAINER_WORKSPACE: "/infmax-workspace" | ||
| AIPERF_TRACE_IDLE_GAP_CAP_SECONDS: "300" |
There was a problem hiding this comment.
🔴 This PR modifies 4 Kimi K3 B200 recipe YAMLs (swapping benchmark.aiperf_server_metrics for AIPERF_SERVER_METRICS_URLS) but does not append a new perf-changelog.yaml entry, violating the AGENTS.md non-negotiable invariant that every recipe addition or modification requires one. Please append a new entry for config-key kimik3-fp4-b200-dynamo-vllm-agentic-dspark with a pr-link to this PR (#2569), describing the AIPerf metrics-URL fix, at the tail of the file.
Extended reasoning...
AGENTS.md § "Non-negotiable benchmark invariants" (line 21) states explicitly: "Every change that can affect benchmark performance and every recipe addition or modification requires a new perf-changelog.yaml entry. The file is append-only and byte-sensitive. Preserve all existing bytes and separator whitespace, and append only at the tail." This PR modifies four recipe YAMLs under benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/, removing the unsupported benchmark.aiperf_server_metrics: true field and adding AIPERF_SERVER_METRICS_URLS to the benchmark.env block. This is unambiguously a "recipe modification" under the stated rule, so a new changelog entry is required regardless of whether the change is purely a config-key fix rather than a topology change.
I confirmed directly against the repository that perf-changelog.yaml (root of repo, ~5820 lines) is not among the PR's 4 changed files, and its tail still ends with two entries for PR #2475 under config-key kimik3-fp4-b200-dynamo-vllm-agentic-dspark — the exact config-key that covers these four recipes. There is no entry referencing PR #2569 (this PR's number) anywhere in the file.
This omission is not merely a style nit: the changelog is the append-only audit trail the repo uses to track exactly which PR touched which recipe and why, and prior modifications to these same recipes were always accompanied by a changelog entry — e.g. PR #2475 (which originally added aiperf_server_metrics: true) appended two entries at what are now lines 5807-5820, and PR #2364 shows a directly analogous precedent, adding a changelog entry specifically for a server-metrics-prefix config change ("Enable SGLang engine metrics and require the sglang: server-metric prefix..."). This PR's change is the vLLM equivalent of that exact kind of metrics-config fix, yet appends nothing.
Step-by-step proof:
- Read AGENTS.md line 21 — invariant requires a changelog entry for every recipe modification, not just performance-affecting ones (the two clauses are joined by "and", not "or", but both apply here as a metrics-related change affects observability of benchmark performance).
- List this PR's changed files — only the 4 recipe YAMLs,
perf-changelog.yamlis absent. tail -40 perf-changelog.yaml— file ends at PR [AgentX] Tune Kimi K3 DSpark on B200 #2475 and [NV] Add H200 DeepSeek-V4-Pro AgentX recipes / [NV] 添加 H200 DeepSeek-V4-Pro AgentX 配方 #2364 entries; no PR Fix Kimi K3 B200 AIPerf metrics configuration #2569 entry exists.grep -n "kimik3-fp4-b200-dynamo-vllm-agentic-dspark" perf-changelog.yaml— shows only the pre-existing [AgentX] Tune Kimi K3 DSpark on B200 #2475 entries covering the recipes this PR touches, confirming these recipes are tracked by the changelog and thus require an update when modified.- Conclusion: the invariant is violated as written.
Fix: Append a new entry at the tail of perf-changelog.yaml (preserving existing bytes/whitespace) with config-keys: [kimik3-fp4-b200-dynamo-vllm-agentic-dspark], a description summarizing the AIPerf metrics-URL fix (replacing the unsupported aiperf_server_metrics field with AIPERF_SERVER_METRICS_URLS env var), and pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2569.
This is flagged as an explicit, non-negotiable repository invariant (not subjective style preference), and the repo's own history shows every prior touch of these recipes carried a changelog entry, so it should be addressed before merge.
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=31558618757 |
基于完整的十九点 AgentX-fast 扫描,将正式扫描收敛到驻留 KV 的 c2/c4/c8/c12,以及 SimpleCPUOffloadConnector 的 c8/c12/c16/c28/c32/c48。保留低延迟曲线、拐点、紧邻的容量边界和高并发对照,同时移除已被支配的重复点。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=31568488368 |
Summary
http://localhost:8000/metricswhile retaining the strictvllm:prefix gate and prompt-token cache reporting.Root cause
PR #2475 was merged while its first fast sweep was still being diagnosed. All 21 jobs in run 31549977914 failed before Slurm submission because the pinned srt-slurm schema rejected
benchmark.aiperf_server_metricsas an unknown field.These recipes use
benchmark.type: custom. The supported wrapper contract isAIPERF_SERVER_METRICS_URLS; it emits exactly one--server-metricsflag followed by the configured URL list.Search result
The corrected 19-point AgentX-fast run 31558618757 completed all throughput jobs and both evals successfully on exact head
258c6e4f544ceff5fdc8d0007b431d69d17b405f.The fast low-latency frontier is resident c2/c4, then offload c8/c12. Resident c12 and offload c16 retain the immediate cliff boundary. Offload c28/c32/c48 remain high-capacity controls because the prior one-hour sweep showed the high-throughput end of the curve is not reliably ranked by the shorter stochastic sample.
Every completed throughput point passed the strict
vllm:server-metrics prefix gate with nonempty artifacts. Realized commands contained exactly one--server-metrics http://localhost:8000/metrics, and AIPerf reported 1/1 endpoint reachable.Validation
srtctl dry-runwith pinned renderer commitdf5baa93f4caf5169dea2a4236ad2cc742fe40e7.git diff --checkand performance changelog validation passed.This PR is English-only.