Skip to content

Add DSV4 B200 disaggregated Dynamo SGLang MTP configuration / 新增 DSV4 B200 分离式 Dynamo SGLang MTP 配置 - #2554

Open
RohitNagraj wants to merge 2 commits into
mainfrom
dsv4-fp4-b200-dynamo-sglang-mtp-nscale
Open

Add DSV4 B200 disaggregated Dynamo SGLang MTP configuration / 新增 DSV4 B200 分离式 Dynamo SGLang MTP 配置#2554
RohitNagraj wants to merge 2 commits into
mainfrom
dsv4-fp4-b200-dynamo-sglang-mtp-nscale

Conversation

@RohitNagraj

@RohitNagraj RohitNagraj commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator

Description

Add a DSV4 B200 disaggregated Dynamo SGLang MTP configuration for 8k/1k.

  • Add five checked-in srt-slurm recipes with EAGLE speculative decoding and chat-formatted benchmark inputs.
  • Use the staged DSV4 checkpoint native quantization metadata and the checkpoint-compatible MoE backend.
  • Use NIXL with UCX RC transport for KV state transfer.
  • Route multi-node runs through the dedicated B200 launcher and pin srt-slurm to the validated revision.

Validation:

  • Generated all five throughput configurations and four evaluation configurations.
  • Validated launcher shell syntax, YAML and srtctl schemas, recipe/config consistency, chat-template usage, and changelog ordering.
  • Passed the matrix-logic test suite.

中文说明

新增 DSV4 在 B200 上的分离式 Dynamo SGLang MTP 8k/1k 配置。

  • 新增五个已检入的 srt-slurm recipe,启用 EAGLE 投机解码和聊天模板格式的基准测试输入。
  • 使用已暂存 DSV4 检查点原生的量化元数据,并选择与检查点兼容的 MoE 后端。
  • 使用 NIXL 和 UCX RC 传输 KV 状态。
  • 通过专用 B200 启动器路由多节点作业,并将 srt-slurm 固定到已验证的版本。

验证:

  • 已成功生成五个吞吐量配置和四个评估配置。
  • 已验证启动器 Shell 语法、YAML 与 srtctl schema、recipe 与配置一致性、聊天模板使用情况以及变更日志顺序。
  • 已通过矩阵逻辑测试套件。

Related Issue

N/A

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /reuse-sweep-run on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

新增 DSV4 B200 nscale UCX 验证配置。
@RohitNagraj
RohitNagraj force-pushed the dsv4-fp4-b200-dynamo-sglang-mtp-nscale branch from 328eff9 to 0555089 Compare August 11, 2026 01:19
@RohitNagraj
RohitNagraj marked this pull request as ready for review August 11, 2026 01:20
@RohitNagraj
RohitNagraj requested a review from a team August 11, 2026 01:20
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@claude

claude Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Claude finished @RohitNagraj's task in 6m 51s —— View job


PR Review — DSV4 FP4 B200 nscale Dynamo-SGLang MTP UCX canary

  • Gather context and read the full diff
  • Review recipe YAML (disagg-b200-mid-curve-1p2d-dep8-dep8-mtp.yaml)
  • Review configs/nvidia-master.yaml changes (image accessibility, config key format)
  • Review perf-changelog.yaml (entry position, present alongside master config change)
  • Review runners/launch_b200-nscale-slurm.sh (enroot import, routing logic)
  • Post review feedback

Summary

No blocking issues found. Two 🟡 warnings posted inline, both on the new recipe: UCX_PROTO_EMULATION_ENABLE does not appear to be a real UCX config variable (the canary's "host-staged protocol emulation" knob may be a no-op — host staging is still forced by UCX_IB_GPU_DIRECT_RDMA=no + cuda_copy, and the already-enabled UCX_PROTO_INFO=y output will show what was actually selected), and cpus-per-task: "144" looks copied from Grace-based GB200/B300 recipes and may be unsatisfiable on the x86 nscale B200 nodes, where no existing recipe sets that directive.

Checks that passed: perf-changelog.yaml updated alongside the master config and appended at the end of the file; image is a public Docker Hub reference matching the recipe's container:; launcher keeps the enroot import docker:// reproducibility pattern and its new framework gate is logically correct (SPEC_DECODING is exported by benchmark-multinode-tmpl.yml); the recipe includes use_chat_template: true for MTP; the dsv4 prefix, dynamo-sglang framework value, and config schema all match established entries; the dynamo hash in the recipe matches the router version in the master config.

Caveat: this sandbox has no network access, so I could not verify externally that the Docker Hub tag nightly-dev-cu13-20260710-cfc66e05 exists, that the pinned srt-slurm commit 04e87fcc is reachable from main (the --single-branch main clone requires it), or the UCX variable question above — the PR's stated manifest/matrix validations presumably cover the first two.


# Keep RC for inter-node transport, but stage CUDA buffers through host memory.
UCX_TLS: "rc,cuda_ipc,cuda_copy,self"
UCX_IB_GPU_DIRECT_RDMA: "no"
UCX_PROTO_EMULATION_ENABLE: "y"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 WARNING: UCX_PROTO_EMULATION_ENABLE does not appear to be a real UCX configuration variable
Why it matters: Mainline UCX exposes UCX_PROTO_ENABLE / UCX_PROTO_INFO, and proto-v2 "emulation" protocols are selected automatically when a transport lacks a capability — I can find no PROTO_EMULATION_ENABLE knob in the UCX config table. Unknown UCX_* vars are silently ignored (at most a "unused env var" warning), so the run will proceed, but the canary's stated third knob ("enable host-staged protocol emulation") would be a no-op and conclusions drawn from this arm could be misattributed. Host staging is still forced here by UCX_IB_GPU_DIRECT_RDMA=no + cuda_copy, so the experiment isn't broken — it just may not test what the PR description claims. (Same variable appears in the decode block at line 77.)
Fix: Verify the variable exists in the container's UCX build (ucx_info -c | grep -i proto), or rely on UCX_PROTO_INFO=y output (already enabled) to confirm which protocol was actually selected, and drop the variable if it's not recognized. Note: I could not verify against UCX source/docs from this sandbox (network access blocked), so please double-check before drawing conclusions from the run.

Comment on lines +17 to +19
sbatch_directives:
cpus-per-task: "144"
mem: "0"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 WARNING: cpus-per-task: "144" looks copied from GB200/B300 recipes and may be unsatisfiable on nscale B200 nodes
Why it matters: 144 matches GB200/GB300 Grace nodes (2×72 cores) — see e.g. disagg-gb200-mid-curve-1p1d-dep8-dep16-mtp.yaml, which this recipe otherwise mirrors. The B200 nscale nodes are x86 (the sibling dsv4-fp4-b200-dynamo-vllm-mtp recipes use the vllm-x86_64 image), and none of the existing B200 nscale recipes set cpus-per-task (they use sbatch_directives: segment: "1" instead, e.g. benchmarks/multi_node/srt-slurm-recipes/vllm/deepseek-v4/8k1k/disagg-b200-1p1d-dep8-dep8-mtp.yaml:15). If the nscale nodes have fewer than 144 cores available per task, sbatch rejects the job with "Requested node configuration is not available" and the canary never starts.
Fix: Confirm the nscale B200 node core count supports --cpus-per-task=144, or drop the directive (and consider whether segment: "1" from the existing nscale recipes is needed) so Slurm uses its defaults.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. This is a "[Do Not Merge]" canary PR whose content is byte-for-byte identical to commit 0555089, already merged to main separately — the same recipe, master-config entry, changelog entry, and launcher gating were already landed and validated. No new risk introduced.

What was reviewed: the new SGLang recipe's UCX/host-staging env vars for prefill vs. decode, the launcher's expanded framework/spec-decoding gate logic in launch_b200-nscale-slurm.sh, and the srt-slurm checkout/recipe-copy branching for the dynamo-sglang path. The flagged missing SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2=1 in prefill_environment was checked against sibling multi-node DSv4 recipes — CAR_V2 is single-node only and correctly omitted here since prefill runs on 1 node with DEP8.

Extended reasoning...

Overview

The PR adds one new benchmark recipe YAML for a DeepSeek-V4-Pro FP4 Dynamo-SGLang MTP run on B200 nscale (1P/2D DEP8, host-staged UCX transport), a corresponding entry in configs/nvidia-master.yaml, a perf-changelog.yaml entry, and an extension of runners/launch_b200-nscale-slurm.sh to route this new framework/spec-decoding combination and pin the srt-slurm revision used for it.

Security risks

None. This is CI/benchmark infrastructure config — no auth, crypto, or user-facing surface. The launcher changes only affect environment-variable gating and which pinned public repo revision is checked out for benchmark execution.

Level of scrutiny

Low-to-moderate is appropriate: this is a benchmark-recipe addition following an extremely well-established pattern in this repo (dozens of similar recipes exist), and it is explicitly marked '[Do Not Merge]' — a canary intended only to validate the config via CI, not to be merged into main as-is. Notably, the entire diff (all 4 files) is byte-for-byte identical to commit 0555089, which already landed on main from the same author. That means this exact change has already been reviewed/merged through a separate path, materially reducing residual risk here.

Other factors

The bug-hunting pass found no issues, and the one candidate concern raised (missing CAR_V2 flag in the new prefill_environment) was checked against sibling recipes and confirmed to be correct as written — CAR_V2 is intentionally single-node-only and prefill here runs on a single node. The launcher's new gating conditional was traced against SPEC_DECODING, which is a standard workflow-exported env var used identically across many other launcher scripts, so no wiring risk there either.

@github-actions

Copy link
Copy Markdown
Contributor

扩展 DSV4 B200 Dynamo-SGLang MTP 配置,补充五个分离式拓扑并保持配置与 recipe 一致。
@RohitNagraj RohitNagraj changed the title [Do Not Merge] Add DSV4 FP4 B200 nscale Dynamo-SGLang MTP UCX canary / 添加 DSV4 FP4 B200 nscale Dynamo-SGLang MTP UCX 验证配置 Add DSV4 B200 disaggregated Dynamo SGLang MTP configuration / 新增 DSV4 B200 分离式 Dynamo SGLang MTP 配置 Aug 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant