Skip to content

[AMD] [AGENTX] MiniMax-M3 FP4 MTP on MI355X - #2487

Open
ajith-sirra-amd wants to merge 27 commits into
mainfrom
amd/agentx-minimax-m3-mtp
Open

[AMD] [AGENTX] MiniMax-M3 FP4 MTP on MI355X#2487
ajith-sirra-amd wants to merge 27 commits into
mainfrom
amd/agentx-minimax-m3-mtp

Conversation

@ajith-sirra-amd

@ajith-sirra-amd ajith-sirra-amd commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator

Add MiniMax-M3 FP4 MI355X Agentic Support with MTP

Adds a new single-node agentic-coding benchmark recipe for MiniMax-M3 (FP4) on MI355X using vLLM, with EAGLE3 MTP speculative decoding and DRAM-backed KV cache offloading (vllm-simple).

Changes

  • New launch script: starts a vLLM server with EAGLE3 speculative decoding, AITER MoE/fusion backends, INT4 quick-reduce all-reduce, FP8 KV cache, TP/EP parallel, and optional native KV offloading to host DRAM, then runs the agentic replay/eval harness.
  • New config entry minimaxm3-fp4-mi355x-vllm-agentic-mtp (image vllm-openai-rocm:nightly-cb8104839c141609d99f1254459ef3a4f1bd4263, tp=4, DRAM KV offloading via vllm-simple, spec-decoding: mtp).
  • Synthetic acceptance for the throughput replay: rejection_sample_method: synthetic with synthetic_acceptance_length: 3.35, the committed MiniMax-M3 EAGLE3 golden AL for thinking_on at num_speculative_tokens=5 (golden_al_distribution/minimaxm3_eagle3.yaml), per the AgentX fairness guidelines. EVAL_ONLY accuracy runs keep real target verification, since synthetic acceptance bypasses verification and corrupts the eval score.
  • Changelog entry documenting the addition (PR link TBD).

Status

Marked [WIP].

@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@ajith-sirra-amd ajith-sirra-amd added the agentx AgentX benchmarks, recipes, and infrastructure label Aug 4, 2026
Comment thread benchmarks/single_node/agentic/minimaxm3_fp4_mi355x_mtp.sh
Comment thread perf-changelog.yaml Outdated
Comment thread benchmarks/single_node/agentic/minimaxm3_fp4_mi355x_mtp.sh Outdated
AjithSirra and others added 5 commits August 4, 2026 17:03
…g Search Space & KV Backend

Signed-off-by: Sirra <asirra@amd.com>
…R Id to Perf Change Log.

Signed-off-by: Sirra <asirra@amd.com>
…RAFT Model & Details.

Signed-off-by: Sirra <asirra@amd.com>
…eptance to the golden AL

Throughput replay now runs with rejection_sample_method=synthetic and
synthetic_acceptance_length=3.35, the committed minimax-m3 EAGLE3 golden AL
for thinking_on at num_speculative_tokens=5
(golden_al_distribution/minimaxm3_eagle3.yaml), as required by the AgentX
fairness guidelines for agentic speculative-decoding submissions.

EVAL_ONLY accuracy runs keep real target verification -- synthetic acceptance
bypasses verification and corrupts the eval score.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Signed-off-by: Sirra <asirra@amd.com>
@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Signed-off-by: Sirra <asirra@amd.com>
Signed-off-by: Sirra <asirra@amd.com>
@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

ajith-sirra-amd and others added 6 commits August 11, 2026 17:04
Preserve the upstream GLM-5.2 and Kimi/DSv4 changelog entries while retaining PR #2487 for correction and expanded performance testing.\n\n合并最新主分支,同时保留上游 GLM-5.2、Kimi/DSv4 更新以及待修正和扩展性能测试的 PR #2487
Restore EAGLE3 with the committed golden acceptance length, add TP4 and TP2 GPU-resident coverage, and add TP-sharded LMCache MP offload using the official ROCm wheel. Require vLLM Prometheus artifacts and preserve a strict no-exception benchmark path.\n\n恢复使用已提交黄金接受长度的 EAGLE3,新增 TP4/TP2 GPU 常驻搜索空间,并通过官方 ROCm wheel 增加按 TP 分片的 LMCache MP 卸载。强制生成 vLLM Prometheus 指标,并保持无例外的严格基准路径。
@cquil11 cquil11 added the agentx-fast Run AgentX throughput with 1 warmup request per lane and a 20-minute profile; not reusable label Aug 11, 2026
显式安装 LMCache MP 所需的 sortedcontainers 运行时依赖,同时保留 vLLM 镜像中经过验证的 torch/ROCm 软件栈。
…m3-tuning

# Conflicts:
#	perf-changelog.yaml
@cquil11 cquil11 added full-sweep-enabled and removed agentx-fast Run AgentX throughput with 1 warmup request per lane and a 20-minute profile; not reusable labels Aug 12, 2026
@cquil11 cquil11 changed the title [AMD] [WIP] [AGENTX] MiniMax-M3 Support on MI355X with MTP [AMD] [AGENTX] MiniMax-M3 FP4 MTP on MI355X Aug 12, 2026
…m3-tuning

# Conflicts:
#	perf-changelog.yaml
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure AMD full-sweep-enabled

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

4 participants