Skip to content

[3/4] [Power] feat: enable official dcgm energy lane for gb200/gb300 1p1d / 打开 gb200/gb300 官方 1P1D 能耗采集 - #2456

Open
edwingao28 wants to merge 6 commits into
wenyao/power-pr3-multinode-consumptionfrom
wenyao/power-pr4-official-energy
Open

[3/4] [Power] feat: enable official dcgm energy lane for gb200/gb300 1p1d / 打开 gb200/gb300 官方 1P1D 能耗采集#2456
edwingao28 wants to merge 6 commits into
wenyao/power-pr3-multinode-consumptionfrom
wenyao/power-pr4-official-energy

Conversation

@edwingao28

@edwingao28 edwingao28 commented Aug 2, 2026

Copy link
Copy Markdown
Collaborator

Merge strategy: this PR merges with a fork pin. The producer pin points at edwingao28/srt-slurm@6fc1bed (the reviewed v2 producer branch; immutable SHA, additionally protected by fork branch wenyao/pr2-pin-validated-v2). The upstream NVIDIA/srt-slurm PR proceeds independently; once it merges, a one-line follow-up PR swaps URL+SHA to upstream and re-smokes both platforms. Pin-swap re-smoke evidence for 6fc1bed: GB200 30969382477 (full ladder, 8/8 power_valid=1, manifest producer==6fc1bed, offline validator exit 0), GB300 30981806370 (full ladder, 8/8 power_valid=1, manifest producer==6fc1bed). The N=3 ladder tables below were measured under the functionally identical v1 pin (6609d46); the v2 runs reproduce them (gb200 c4 5.180 vs N=3 mean 5.219; gb300 c4 4.586 vs 4.593, c128 1.011 vs 1.010).

中文:本 PR 以 fork pin 合入(SHA 内容寻址不可变 + 兜底保护分支),上游 PR 独立推进,合入后一个 one-line 跟进 PR 换成 upstream SHA 并双平台 re-smoke。

Stacked on #2437 (multinode consumer), which stacks on #2323 (single-node).

What this adds

  • telemetry block in the two official 1P1D 8k1k recipes. gb300 uses port 19401: 9401 is already bound by the cluster-level exporter on im-gb300 nodes.
  • lane-scoped provisioning in the gb200/gb300 launchers: a run opts in only when its recipe carries an enabled dcgm-power block. Every other model/recipe keeps the exact same path as before.
  • producer pin contract: power lanes clone the pinned SHA, assert HEAD matches, and write power-producer-sha.txt. CI derives POWER_PRODUCER_SHA from that stamp; the workflow input stays as a manual override.
  • contract tests for the above (utils/test_gb200_power_official_contract.py, utils/test_gb300_power_official_contract.py).

Scope: 1P1D only. No nvidia-master.yaml change. Dispatches for this PR use exact-key test-config, not full-sweep.

Validation so far (fork pin): GB200 run 30618706258 (340.5 W/GPU avg, 193,404 J), GB300 run 30663050396 (349.3 W/GPU avg, 172,800 J), both with power_valid=1 and stored==recomputed sidecars. Local: 124 tests pass.

中文:为 gb200/gb300 官方 1P1D lane 打开 DCGM 能耗采集。recipe 声明 telemetry(gb300 用 19401 端口),launcher 按 recipe 判定是否 provision exporter 与 pin producer(非 power lane 路径不变),CI 从 launcher stamp 读取 POWER_PRODUCER_SHA。当前 pin 指向已验证的 fork SHA,upstream merge 后换 pin 并重跑双平台 c4 smoke,再转正式 review。


Official ladder results (D1 runs on this branch, fork pin; one dispatch per platform runs the recipe's full concurrency sweep): GB200 30951410151, GB300 30951412632. All 16 windows power_valid=1, manifest producer SHA == pin, stored == recomputed sidecars, per-GPU sample gaps ≤ 1.03 s.

hw conc power_valid W/GPU total W total J prefill J decode J J/1k out-tok J/1k in-tok out tok/s/GPU mean TTFT s mean TPOT s
gb200 1 1 283.0 2,263.8 131,153.9 54,320.9 76,833.0 14,098.03 1,772.52 40.1 0.339 0.0059
gb200 2 1 310.8 2,486.8 147,494.7 59,213.5 88,281.1 7,885.31 996.42 78.8 0.366 0.0059
gb200 4 1 347.1 2,777.1 189,752.8 74,233.4 115,519.4 5,155.63 645.13 134.7 0.538 0.0066
gb200 8 1 405.2 3,241.2 258,333.2 98,373.9 159,959.3 3,487.74 441.62 232.3 0.634 0.0077
gb200 16 1 472.2 3,777.8 362,887.9 141,556.9 221,331.0 2,478.30 308.75 381.1 0.768 0.0093
gb200 32 1 557.3 4,458.0 550,756.1 220,795.0 329,961.1 1,859.76 234.09 599.3 1.042 0.0118
gb200 64 1 662.8 5,302.5 827,651.4 359,795.6 467,855.8 1,402.97 175.02 944.9 1.355 0.0149
gb200 128 1 786.6 6,292.5 1,273,182.6 614,364.1 658,818.6 1,081.18 134.63 1,455.0 2.163 0.0190
gb300 1 1 289.8 2,318.5 113,820.3 47,056.1 66,764.3 12,234.80 1,538.26 47.4 0.352 0.0049
gb300 2 1 313.5 2,508.1 133,353.4 53,740.4 79,612.9 7,129.29 900.89 88.0 0.513 0.0051
gb300 4 1 354.1 2,833.1 169,519.3 66,519.6 102,999.7 4,605.88 576.34 153.8 0.463 0.0058
gb300 8 1 410.3 3,282.4 233,869.5 90,374.3 143,495.2 3,157.45 399.80 259.9 0.586 0.0068
gb300 16 1 477.9 3,823.5 332,993.4 129,314.6 203,678.9 2,274.14 283.31 420.3 0.751 0.0084
gb300 32 1 563.9 4,511.5 513,995.9 205,957.3 308,038.6 1,735.63 218.47 649.8 0.982 0.0109
gb300 64 1 663.1 5,305.2 768,153.2 334,258.7 433,894.5 1,302.12 162.44 1,018.6 1.233 0.0139
gb300 128 1 780.9 6,247.0 1,191,762.0 578,366.1 613,395.9 1,012.03 126.02 1,543.2 2.192 0.0177

中文:上表为两平台正式 ladder(每平台一次 dispatch 覆盖 c1–128 全部 8 窗)。GB300 每输出 token 能耗全程低于 GB200(c4 低 ~10.7%,c128 低 ~6.4%),且 TPOT 同步更优;两平台到 c128 尚未出现能效饱和。


Reproducibility, N=3 per platform (mean over 3 independent dispatches; CV% in parentheses). GB200: 30951410151, 30960242894, 30962794356. GB300: 30951412632, 30958592389, 30958596147. All 48 windows power_valid=1. Worst CV across the 16 (platform, conc) cells: 1.59% on J/out-tok, 1.10% on W/GPU.

hw conc N J/out-tok (CV%) W/GPU (CV%) J/query out tok/s/GPU TTFT s TPOT s
gb200 1 3 14.162 (0.77) 282.3 (0.28) 13,174.9 39.9 0.346 0.0059
gb200 2 3 7.905 (0.24) 310.5 (0.10) 7,393.3 78.6 0.368 0.0059
gb200 4 3 5.219 (1.59) 346.9 (0.06) 4,802.3 133.0 0.544 0.0067
gb200 8 3 3.497 (0.24) 406.4 (0.31) 3,237.4 232.4 0.631 0.0077
gb200 16 3 2.477 (0.16) 470.9 (0.25) 2,266.7 380.2 0.794 0.0093
gb200 32 3 1.864 (0.36) 557.2 (0.10) 1,724.6 598.0 1.067 0.0118
gb200 64 3 1.402 (0.35) 663.1 (0.05) 1,292.8 945.6 1.359 0.0149
gb200 128 3 1.081 (0.28) 787.2 (0.16) 994.1 1,457.1 2.100 0.0190
gb300 1 3 12.282 (0.67) 289.1 (1.10) 11,426.4 47.1 0.412 0.0049
gb300 2 3 7.104 (0.60) 313.1 (0.78) 6,644.2 88.2 0.480 0.0051
gb300 4 3 4.593 (0.61) 352.8 (0.73) 4,225.8 153.6 0.460 0.0058
gb300 8 3 3.165 (0.58) 408.9 (0.54) 2,930.7 258.3 0.611 0.0068
gb300 16 3 2.274 (0.06) 477.7 (0.07) 2,081.3 420.1 0.735 0.0084
gb300 32 3 1.728 (0.41) 562.0 (0.30) 1,598.7 650.7 0.967 0.0109
gb300 64 3 1.302 (0.27) 661.9 (0.18) 1,199.7 1,017.0 1.252 0.0139
gb300 128 3 1.010 (0.26) 782.8 (0.37) 929.3 1,549.9 2.098 0.0177

中文:每平台 3 次独立 dispatch(重新部署、重新采样)后按窗口聚合。16 个 (平台, 并发) 单元格里 J/out-tok 的 CV 最差 1.59%(gb200 c4),其余全部 ≤0.77%;W/GPU 最差 1.10%。仪器与结果均可复现。两次 gb200 失败 run(30958590082/30958594280)为同集群并发 dispatch 导致的冷启动互踩、非仪器问题,已剔除并以串行重跑替代。

@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@edwingao28
edwingao28 marked this pull request as ready for review August 3, 2026 06:25
@claude

claude Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Claude finished @edwingao28's task in 3m 59s —— View job


Review — PR #2456 (official dcgm energy lane, gb200/gb300 1P1D)

  • Gather context (diff vs wenyao/power-pr3-multinode-consumption, read changed files)
  • Review launcher lane-gating & producer-pin logic
  • Review recipe telemetry blocks & workflow changes
  • Verify contract tests (statically — pytest execution isn't permitted in this environment; CI runs them via test-process-result.yml)
  • Post review feedback

LGTM - no blocking issues found.

Things I specifically checked and confirmed sound:

  • Lane detection ordering: USES_DCGM_POWER is derived from the workspace recipe mirror before the srt-slurm clone in both launchers, and CONFIG_FILE's :selector suffix is stripped (${CONFIG_FILE%%:*}) before path resolution, matching how CONFIG_PATH is handled later.
  • Non-power path preservation: every power addition (exporter import, pinned clone, sed yaml injection, provenance copies) is gated on USES_DCGM_POWER == 1; gb300's fp8/fp4 ref logic (v1.0.25 / sa-submission-q2-2026) is byte-preserved in the else branch.
  • sed injection target: nginx-sqsh: exists exactly once in both generated srtslurm.yaml heredocs, so the append lands inside the containers map in both launchers.
  • Workflow wiring: POWER_PRODUCER_SHA is job-level env from the input, so the -z check correctly treats the input as an override; the stamp is written to $GITHUB_WORKSPACE during the launch step and read relative to the default working dir in the same job. Stale-stamp carryover across runs is prevented by actions/checkout with clean: true before the launch step.
  • New repo-wide invariant (test_workflows_carry_no_producer_sha_literal): passes against the current tree (no non-uses: 40-hex literals in any workflow).

One non-blocking observation (no action needed now): the lane check is two independent greps (provider: dcgm-power + enabled: true anywhere in the recipe), so a future recipe with telemetry.enabled: false plus an unrelated enabled: true elsewhere would false-positive. The test_exactly_two_recipes_opt_into_dcgm_power contract currently fences this in, so it's fine for this PR's scope — just something to keep in mind if telemetry blocks spread to more recipes.

Also noting the PR's own plan: the producer pin currently points at the fork SHA (edwingao28/srt-slurm@6609d46) — per the description this swaps to the upstream merge SHA plus a two-platform c4 smoke before merge. The test_pin_constants_identical_across_launchers / test_pin_literal_lives_only_in_the_two_launchers tests will keep that swap honest.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No bugs were found, but I'm deferring to a human reviewer given the scope and risk profile of this change. This PR is explicitly marked draft-on-purpose — the producer pin currently points at a third-party fork (edwingao28/srt-slurm) rather than the eventual upstream merge SHA — and it modifies the critical gb200/gb300 multinode launcher scripts that gate all benchmark submissions on those clusters, not just the new power lane.

What was reviewed: the recipe-driven power-lane detection logic in both launchers (grep-based, resolved against the workspace recipe mirror before the srt-slurm clone exists), the pinned-clone-and-assert-HEAD contract that writes power-producer-sha.txt, the CI step in benchmark-multinode-tmpl.yml that derives POWER_PRODUCER_SHA from that stamp (input still overrides), and the new contract tests covering both launchers. The bug hunting system additionally checked the gb300 dcgm-exporter image enroot import syntax and ruled it out as correct.

Extended reasoning...

Overview

This PR wires an official DCGM power-telemetry lane into the gb200/gb300 1P1D qwen3.5 recipes. It adds a telemetry block to the two recipe YAMLs, adds recipe-gated (grep-detected) exporter provisioning and a pinned srt-slurm producer clone to both runners/launch_gb200-nv.sh and runners/launch_gb300-nv.sh, threads a power-producer-sha.txt stamp from the launcher through benchmark-multinode-tmpl.yml into POWER_PRODUCER_SHA, and adds static contract tests (utils/test_gb200_power_official_contract.py, utils/test_gb300_power_official_contract.py) asserting the launcher text matches the intended contract. Non-power lanes are intended to be byte-identical to their pre-PR behavior (exporter/pin logic is gated behind a USES_DCGM_POWER flag derived from the recipe).

Security risks

The main risk is supply-chain: the power lane clones and executes setup/build steps from a personal fork (edwingao28/srt-slurm) rather than the org-controlled NVIDIA/srt-slurm. This is mitigated by pinning to an exact 40-hex commit SHA and asserting git rev-parse HEAD matches it before proceeding, and the PR description itself states this is a temporary, draft-only state pending the real upstream merge SHA. Still, until that swap happens, CI on official gb200/gb300 power-lane runs executes code from a non-org-controlled repository. No other injection/auth-bypass/data-exposure concerns stood out — the new grep-based recipe detection only reads local repo files, and the sed/yaml edits are scoped to generated config, not user-controlled input.

Level of scrutiny

This warrants a higher level of scrutiny than a typical recipe or config PR: it touches the shared launcher scripts (runners/launch_gb200-nv.sh, runners/launch_gb300-nv.sh) that every model/precision/framework combination on those clusters depends on, and it is explicitly flagged by the author as draft/not-yet-final (fork pin instead of upstream pin, pending a second smoke test on both platforms). A human familiar with the srt-slurm producer and the gb200/gb300 CI paths should confirm the pin-swap plan and re-validate before this leaves draft.

Other factors

The author has already done meaningful validation (real GB200/GB300 runs with power_valid=1 and stored==recomputed sidecars, 124 local tests passing) and added static contract tests enforcing the launcher/recipe/workflow text stays in sync. The bug-hunting system found no bugs. Given the explicit draft status and the fork-pin caveat called out by the author themselves, this PR is not yet intended to be merged as-is, so deferring for human sign-off rather than approving is the right call.

@edwingao28 edwingao28 changed the title feat(power): enable official dcgm energy lane for gb200/gb300 1p1d / 打开 gb200/gb300 官方 1P1D 能耗采集 [3/4] [Power] feat: enable official dcgm energy lane for gb200/gb300 1p1d / 打开 gb200/gb300 官方 1P1D 能耗采集 Aug 4, 2026
@edwingao28
edwingao28 requested a review from a team August 5, 2026 05:50
@edwingao28
edwingao28 force-pushed the wenyao/power-pr4-official-energy branch from 43d76f6 to 2daca89 Compare August 5, 2026 05:50
…in gb launchers

中文:launcher 按 recipe 判定 power lane,provision exporter 并 pin producer;非 power lane 行为不变。
…de template

中文:CI 从 launcher stamp 读取 producer SHA,workflow input 保留为手动覆盖。
中文:两个官方 1P1D recipe 声明 telemetry;gb300 用 19401 端口避开集群级 exporter。
中文:契约测试覆盖 lane 检测、恰好两个 power recipe、pin 单一来源与 stamp 一致性。
@edwingao28
edwingao28 force-pushed the wenyao/power-pr4-official-energy branch from 86c65bb to 048e4e4 Compare August 5, 2026 06:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant