MSCCLPP EP implementation - #852
Draft
Binyang Li (Binyang2014) wants to merge 141 commits into
Draft
Conversation
Port DeepEP's high-throughput MoE dispatch/combine kernels onto MSCCL++
as an optional build target `mscclpp_ep_cpp`, gated by -DMSCCLPP_BUILD_EXT_EP
(OFF by default). Sources are lifted from DeepEP branch
`chhwang/dev-atomic-add-cleanup` and rebased onto upstream MSCCL++ APIs;
the NVSHMEM / IBGDA dependencies are replaced with `PortChannel` +
`MemoryChannel` + the new `Connection::atomicAdd` primitive.
Scope
-----
Intranode (NVLink-only):
* `Buffer` ctor/dtor: cudaMalloc nvl workspace, export IPC handle,
allocate FIFO + peer-pointer tables, start `ProxyService`.
* `sync()`: import peer IPC handles, upload peer pointer table,
build `MemoryDevice2DeviceSemaphore` + `MemoryChannel` per peer.
* `get_dispatch_layout`, `intranode_dispatch`, `intranode_combine`
ported verbatim (torch::Tensor ABI preserved).
Internode HT (NVLink + RDMA):
* `sync()` RDMA branch: cudaMalloc RDMA buffer + `bootstrap->barrier()`
(replacing NVSHMEM symmetric-heap allocation); register with
`all_transport`, exchange via `sendMemory`/`recvMemory`, build 12 IB
QPs/peer + 16 semaphores/peer + 16 port channels/peer.
* Full `internode.cu` port (notify_dispatch / dispatch / cached_notify
/ combine / get_dispatch_layout). The 4 raw `ChannelTrigger` atomic
sites are rewritten to call the new
`PortChannelDeviceHandle::atomicAdd(offset, value)` API; the single
`nvshmem_fence()` is replaced with `__threadfence_system()` (remote
visibility guaranteed by the subsequent port-channel barrier).
* `internode_dispatch` / `internode_combine` host code ported, with
the torch tensor marshalling and CPU spin-wait on mapped counters.
Low-latency (pure RDMA):
* Not ported. `low_latency_dispatch`, `low_latency_combine`,
`clean_low_latency_buffer`, `get_next_low_latency_combine_buffer`
throw `std::runtime_error`; the Python frontend refuses to
construct a Buffer with `low_latency_mode=True`.
Python layer
------------
* New pybind11 + libtorch Python extension `mscclpp_ep_cpp` (separate
from the nanobind `_mscclpp` because the EP ABI carries
`torch::Tensor` / `at::cuda::CUDAStream`).
* `mscclpp.ext.ep.Buffer` mirrors `deep_ep.Buffer`; exchanges device
IDs, IPC handles and the bootstrap UniqueId over the user's
`torch.distributed` process group before calling `sync()`.
* `mscclpp.ext` auto-imports `ep` if the extension is built.
Build
-----
* `src/ext/ep/CMakeLists.txt`: finds Python + Torch; warns and skips if
`CMAKE_PREFIX_PATH` doesn't point at `torch.utils.cmake_prefix_path`.
Falls back to Torch's bundled pybind11 if a standalone pybind11 is not
installed. Links `libtorch_python` explicitly (without it, `import
mscclpp_ep_cpp` fails with `undefined symbol: THPDtypeType`).
* Top-level `CMakeLists.txt` exposes the `MSCCLPP_BUILD_EXT_EP` option
(default OFF).
Tests
-----
* `test/python/ext/ep/test_ep_smoke.py`: skipped if the extension isn't
built. Covers Config round-trip, low-latency size hint, and the LL
construction guard. Multi-rank functional tests still to do on H100.
Notes
-----
* Builds against the preceding "atomic add" commit which adds
`Connection::atomicAdd` and `PortChannelDeviceHandle::atomicAdd` to
upstream MSCCL++.
* Intranode path verified end-to-end (build + import + smoke tests).
* Internode HT is code-complete but requires real IB hardware to
validate; see `src/ext/ep/README.md` for the detailed port plan and
remaining LL migration.
Port DeepEP's pure-RDMA low-latency (LL) MoE kernels from
csrc/kernels/internode_ll.cu (branch chhwang/dev-atomic-add-cleanup)
into the MSCCL++ EP extension. NVSHMEM / IBGDA device primitives are
replaced with MSCCL++ PortChannelDeviceHandle operations:
nvshmemx_barrier_all_block() -> port-channel signal+wait ring
nvshmemi_ibgda_put_nbi_warp(...) -> lane-0 PortChannel.put(...)
nvshmemi_ibgda_amo_nonfetch_add(...) -> lane-0 PortChannel.atomicAdd(...)
The atomicAdd path relies on the MSCCL++ Connection::atomicAdd /
PortChannelDeviceHandle::atomicAdd API cherry-picked from branch
chhwang/new-atomic-add; the LL dispatch path uses a signed delta
(-num_tokens_sent - 1) which the new int64_t signature supports.
Changes:
* New file src/ext/ep/kernels/internode_ll.cu (~530 lines) with the
three kernels clean_low_latency_buffer, dispatch<kUseFP8,...>,
combine<...> plus their launchers. rdma_buffer_ptr is threaded
through the launchers so the kernel can translate virtual addresses
into registered-memory offsets expected by MSCCL++.
* kernels/api.cuh: replace the single stub signature with full LL
launcher prototypes.
* buffer.cc: replace the four LL throw-stubs
(clean_low_latency_buffer, low_latency_dispatch,
low_latency_combine, get_next_low_latency_combine_buffer) with
torch-Tensor implementations ported from DeepEP/csrc/deep_ep.cpp.
* Drop src/ext/ep/internode_stub.cc and its CMake entry.
* python/mscclpp/ext/ep/buffer.py: remove the low_latency_mode=True
NotImplementedError guard; update docstring.
* test/python/ext/ep/test_ep_smoke.py: rename
test_low_latency_rejected -> test_low_latency_buffer_construct
to reflect that LL construction is now accepted.
* src/ext/ep/README.md: update status matrix, document the
NVSHMEM -> MSCCL++ translation table, and list the known
limitations.
This is a structural port: the kernels compile, link, and pass the
single-rank smoke tests, but end-to-end behaviour on multi-node H100
is not yet validated. Two known caveats:
1. Performance will NOT match IBGDA because MSCCL++ port channels
use a CPU proxy; this port is for functional parity, not latency.
2. Buffer::sync() in LL mode only connects peers that share the
same local GPU id (DeepEP convention), so the LL kernels assume
a one-GPU-per-node topology (num_ranks == num_rdma_ranks).
Multi-GPU-per-node LL layouts will need a follow-up in sync().
Tested:
cmake --build build -j --target mscclpp_ep_cpp # builds clean
pytest test/python/ext/ep/test_ep_smoke.py # 3 passed
Three issues blocked end-to-end intranode validation across multiple ranks. This commit fixes them and adds a 2/4/8-rank functional test. 1. Combine receiver: OOB __shared__ read In the combine receiver warp, the wait loop evaluated `channel_tail_idx[recv_lane_id] <= expected_head` before the `expected_head >= 0` guard. `channel_tail_idx` is a shared array of size `kNumRanks`, but the loop runs on all 32 lanes of a warp, so lanes with `recv_lane_id >= kNumRanks` indexed out of bounds. compute-sanitizer reported "Invalid __shared__ read of size 4 bytes" at combine<bf16,2,768>+0xdd0, surfaced asynchronously as cudaErrorIllegalAddress at the kernel launch site. Swap the operands so the rank-bounds check short-circuits the shared read. 2. Python bindings: UniqueId ABI `mscclpp::UniqueId` is a `std::array<uint8_t, N>` which pybind11 auto-converts to a Python `list`, silently overriding any `py::class_<UniqueId>` wrapper. Expose `create_unique_id` / `connect` as lambdas that produce/consume `py::bytes` and memcpy into a local `UniqueId`. Also coerce `bytes`->`bytearray` at the Python call site for `sync()` whose signature expects `pybind11::bytearray`. 3. Python frontend: communicator required for NVL-only sync `Buffer::sync()` uses `communicator->connect(ipc_config, ...)` on the pure-NVLink path, so the communicator must be initialized even when `num_rdma_ranks == 1` and `low_latency_mode == False`. Always broadcast the unique id and call `runtime.connect()` before `sync()`. Validation on a single H100x8 node via torchrun: - 2 ranks: dispatch 195 tokens, combine diff=0 - 4 ranks: dispatch 371 tokens, combine diff=0 - 8 ranks: dispatch 456 tokens, combine diff=0 Test harness added at test/python/ext/ep/test_intranode_multirank.py.
The `internode` kernels index device-side port channel handles as
`port_channel_handles[channel_id * num_ranks + peer_rank]`, where
`peer_rank` is a global rank in [0, num_ranks). `Buffer::sync` was
building that table by iterating `std::unordered_map<int, MemoryId>`
(and similarly for connections/semaphores), which yields hash order
rather than ascending rank order. Once the cross-node fan-out grew
beyond a single peer, a local rank's trigger for peer `r` landed on
the semaphore/memory pair of a different peer, so RDMA puts and
atomic tail updates went to the wrong destination and the forwarder
spun on a tail counter that never advanced.
Changes:
- Build `sema_ids` and `port_channel_handles` by iterating
`for (int r = 0; r < num_ranks; ++r)` and looking up the
connection / memory id for rank `r`, skipping ranks excluded by
low-latency mode (inserting a placeholder handle so the stride
stays `num_ranks`).
- Tag the RDMA-phase `sendMemory`/`recvMemory`/`connect` calls with
`kRdmaTag = 1` so they do not collide with NVL-phase tag-0
traffic between the same pair of ranks.
- Drop an unused `r` local in the NVL setup loop.
With this fix and a matched `libmscclpp.so` on both nodes, the
2-node x 8-GPU internode HT dispatch path completes successfully
(`[dispatch] OK`). Combine is still under investigation.
Also adds `test/python/ext/ep/test_internode_multirank.py`, a
torchrun-based 2-node functional test that exercises
`get_dispatch_layout` -> `internode_dispatch` -> `internode_combine`
and validates per-source-rank token values end-to-end.
Two issues prevented internode HT combine from completing on 2x8 H100: 1. Wrong prefix matrices passed to internode_combine. Combine runs in the reverse direction of dispatch, so it must consume the receiver-side matrices returned by dispatch (recv_rdma_channel_prefix_matrix, recv_rdma_rank_prefix_sum, recv_gbl_channel_prefix_matrix), not the sender-side rdma_channel_prefix_matrix / gbl_channel_prefix_matrix. This matches DeepEP's deep_ep/buffer.py::internode_combine handle unpacking. Without the fix the NVL forwarder's 'NVL check' timed out because token_start_idx/token_end_idx were computed against the wrong per-channel layout. 2. Cross-rank race between dispatch and combine. Even with the correct matrices, launching combine immediately after dispatch deadlocked the forwarder NVL check (tail stuck one short of expected_head) because peers still had in-flight dispatch proxy traffic while fast ranks had already started combine. A torch.cuda.synchronize() + dist.barrier() between the two calls makes the test pass deterministically on 16 ranks (combine diff == 0, max|expected| up to 60.0). The barrier in the test is a workaround; the real fix belongs in Buffer::internode_dispatch / Buffer::internode_combine so the dispatch->combine handoff fully fences outstanding proxy work across ranks. Marked with an XXX comment in the test.
Refresh status docs and comments now that internode HT dispatch and combine have been validated end-to-end on 2 nodes x 8 H100 GPUs via test/python/ext/ep/test_internode_multirank.py (all 16 ranks recover their per-rank token payloads with zero diff). - src/ext/ep/README.md: consolidate the previously duplicated README into a single document; mark intranode and internode HT dispatch and combine as validated in the status table; add a 'Running the tests' section with torchrun examples for both the intranode and the 2x8 internode setups; record the dispatch->combine torch.cuda.synchronize() + dist.barrier() requirement under Known limitations; mark Phase 2 DONE and keep Phase 3 (LL) as structural port, untested. - python/mscclpp/ext/ep/buffer.py: update the module docstring and the Buffer constructor docstring to say internode HT is validated and clarify that LL mode is untested on multi-node hardware. - src/ext/ep/buffer.cc: drop the stale 'NVSHMEM support not yet ported' and 'low-latency paths still stubbed' comments. mscclpp_ep does not use NVSHMEM at all (PortChannel/MemoryChannel replace it), and the LL paths are a structural port that is present but untested, not stubbed. Note validation on 2x H100x8 in the internode section header.
- Buffer::sync no longer drops non-same-GPU-id peers in low_latency_mode. DeepEP's original filter was safe because its LL path used NVSHMEM; this port drives LL via PortChannel so the kernel indexes port_channel_handles[local_expert*num_ranks + dst_rank] for every dst_rank. All peers now get a real memory/connection/semaphore/port channel entry. - Add test/python/ext/ep/test_low_latency_multirank.py (LL dispatch+combine functional round-trip, BF16 only). Works cross-node in DeepEP's 1-GPU-per-node topology. - Known limitation documented in src/ext/ep/README.md and the test docstring: intra-node 8-GPU LL currently hangs because every peer transfer routes through the CPU proxy over IB loopback between distinct HCAs on the same host, and (separately) CudaIpcConnection::atomicAdd is a 64-bit op which mis-aligns the 32-bit rdma_recv_count slots when used for same-node peers. Proper fix needs a mixed-transport LL variant (MemoryChannel for same-node, PortChannel for cross-node) or 64-bit counters.
Gated behind MSCCLPP_EP_BENCH=1 to keep correctness runs fast. Reports per-iter latency (max across ranks, CUDA-event timed) and aggregate effective bandwidth (sum across ranks, dispatch+combine payload bytes). Tunable via MSCCLPP_EP_BENCH_WARMUP / _ITERS / _TOKENS / _HIDDEN. Bench reuses the Buffer allocated for the correctness phase and self-skips if the requested hidden exceeds the per-peer NVL/RDMA budget.
Previously the optional benchmark measured full round-trip latency. Split it to time dispatch alone (N iters) and combine alone (N iters reusing one dispatch output), reporting per-phase latency (max across ranks) and aggregate effective bandwidth (sum across ranks). Applies to intranode HT, internode HT, and the (currently unreachable on intra-node 8-GPU) LL test. Internode HT keeps the sync+barrier guard between dispatch and combine but excludes it from either phase's timing.
…o int64 The low-latency dispatch/combine kernels signal recv counts via MSCCL++ PortChannel.atomicAdd, which lowers to IB IBV_WR_ATOMIC_FETCH_AND_ADD. That opcode requires the remote address to be 8-byte aligned, but LowLatencyLayout packed the per-expert signaling slots as int32. Odd slots landed at offset %8 == 4; the NIC silently dropped those atomics and the target rank spun forever in recv_hook (observed: even->odd direction works, odd->even does not, across all tested topologies including 2-rank intra-node, 8-rank intra-node, and 2-node 1-GPU-each). Widen dispatch_rdma_recv_count_buffer / combine_rdma_recv_flag_buffer to int64_t, update clean kernel + kernel signatures + next_clean pointers accordingly, and add int64_t overloads for st_na_release / ld_acquire_sys_global in utils.cuh. Also drop the bogus self CUDA-IPC connection in Buffer::sync() that was previously skewing the cross-rank buildAndAddSemaphore handshake order; the kernel's same-rank branch uses a direct warp copy and never touches the self port-channel slot (filled with a zero-initialized placeholder so the [local_expert*num_ranks + dst_rank] indexing still holds).
Dropping the self ipc_cfg connection caused cudaErrorInvalidResourceHandle on multi-node launches. Keep the self connection (needed by other code paths that assume every rank is in the connections map) but continue to skip the self slot in the semaphore + port-channel construction loops so the kernel's [local_expert*num_ranks + dst_rank] indexing hits only peer handles; the self slot is a zero-initialized placeholder since the kernel's same-rank branch uses a direct warp copy.
The prior commit skipped r==rank in the semaphore and port-channel build loops on the theory that the self-slot handshake skew was the cause of LL direction asymmetry. That was wrong (the real bug was int32 atomic alignment), and skipping self breaks other code paths that assume every rank slot is represented -- cross-node HT and LL failed with cudaErrorInvalidResourceHandle at the first barrier after Buffer init. Restore the self-inclusive loop.
When all ranks live on the same host (num_rdma_ranks == 1), the LL
kernels now bypass PortChannel/IB-loopback entirely. In Buffer::sync()
we additionally:
- allGather IPC handles for each rank's rdma_buffer_ptr and
cudaIpcOpenMemHandle them into peer_rdma_bases[]
- build per-peer MemoryChannels over CUDA IPC connections (tag=2)
used only for the LL barrier ring
The three LL kernels (clean / dispatch / combine) gain a kIpcPath
template parameter and two extra args (peer_rdma_bases,
memory_channel_handles). At each peer op:
- put -> peer-mapped warp copy over NVLink
- atomicAdd-like flag store -> single-writer st_na_release on peer ptr
- signal/wait barrier -> MemoryChannel signal/wait
Cross-node LL (num_rdma_ranks > 1) is untouched; the IPC setup block is
a no-op. The host launch wrappers select the variant via use_ipc_path.
Each local expert sends one copy per dispatched token back to its owner, so the bytes actually on the wire during combine match dispatch. The previous num_tokens×hidden under-counted by ~num_topk×, making combine BW look artificially low next to dispatch.
- Report both per-rank and aggregate BW to align with NCCL-EP's ep_bench (which reports per-rank GB/s). - Accept MSCCLPP_EP_LL_TOKENS/HIDDEN/TOPK/EXPERTS_PER_RANK env overrides so we can match external benchmark problem sizes (NCCL-EP LL defaults are num_tokens=128, hidden=7168, top_k=8).
Same alignment with NCCL-EP ep_bench as the LL test: report both per-rank (agg/num_ranks) and aggregate throughput.
LL dispatch/combine are latency-bound at typical problem sizes: for num_experts=32 the previous grid was cell_div(32,3)=11 blocks, i.e. 8% of a 132-SM H100. The recv-side bodies already stride tokens by sm_id, so extra blocks parallelize token work linearly. Extra blocks past num_experts are gated out of the send/count phases by the existing 'responsible_expert_idx < num_experts' check. Cap at the device's SM count (cooperative launch + launch_bounds(960,1) allow one block per SM).
On the PortChannel (cross-node) path the extra blocks don't help: the dispatch recv loop strides tokens per-warp-group (not per-SM), and the additional blocks instead add cooperative-grid sync overhead and increase concurrent host-proxy FIFO traffic. Measured cross-node dispatch regressed from 1013us to 3063us when the unconditional grid bump was active. Keep the scaled grid for the IPC path (intra-node), where combine-recv and dispatch token striding scale with sm_id and the 1.2-1.3x speedup reproduces.
The LL combine benchmark was cloning the ~58 MB dispatch recv buffer
('recv_x.clone()') on every timed iteration, adding ~20 us of D2D
memcpy per sample and masking kernel-level changes. It also called
torch.empty() for the output inside the loop. Both now live outside
the timed region; the kernel is invoked against a persistent bench_out
and the recv_x produced by the most recent dispatch.
NCCL-EP's LL dispatch/combine kernel uses (numWarpGroups=1,
numWarpsPerGroup=32) when num_experts <= device_num_sms, giving each
SM ownership of a single expert and 32 warps to cooperate on its
recv-side per-(expert, src_rank) work. We were using (3, 10) — 3
experts per SM, 10 warps per (expert, rank) pair — which left a
significant amount of recv-side parallelism on the table because each
warp had to walk ~3x more tokens sequentially.
Switching to (1, 32) for both dispatch and combine matches NCCL-EP's
structure for typical EP sizes (num_experts in {32, 64, 256}) where
num_experts <= 132 SMs.
The static_assert kNumMaxTopK + 1 <= kNumWarpGroups * kNumWarpsPerGroup
still holds (9 <= 32) and the wider block also lets the staging loop
process the hidden-dim with one int4 per thread (hidden_bf16_int4=896
fits easily in 992 working threads).
Cross-node LL regressed when (1, 32) was applied uniformly: dispatch 1031us -> 1570us, combine 2553us -> 3484us. Larger grid means more concurrent putWithSignal calls onto the host-proxy FIFO and a costlier cg::this_grid().sync() between phases, both of which dominate the IB path even though more SMs help the recv-side compute. Make (kNumWarpGroups, kNumWarpsPerGroup) path-dependent: (1, 32) when use_ipc_path, (3, 10) otherwise. Restores cross-node performance and keeps the intra-node win.
- Add MSCCLPP_EP_BENCH_EXPERTS / _TOPK env knobs so the bench phase can match NCCL-EP's `ep_bench -a ht` defaults (256 experts, top-8). The functional check above continues to use the smaller (num_ranks*4 experts, topk=4) configuration. - Switch BW accounting from recv_tokens*hidden to bench_tokens*hidden, matching NCCL-EP's `RDMA_send` per-rank byte count. The previous formula counted DeepEP's expanded recv layout (one row per (token,src_rank) pair), inflating reported GB/s ~5x and making cross-stack comparisons misleading.
Same change as the intra-node bench (commit 4ed6f22), applied to the cross-node test: - Add MSCCLPP_EP_BENCH_EXPERTS / _TOPK env knobs so the bench phase can match NCCL-EP's `ep_bench -a ht` defaults (256 experts, top-8). - Switch BW accounting from recv_tokens*hidden to bench_tokens*hidden, matching NCCL-EP's `RDMA_send` per-rank byte count.
Each mscclpp::ProxyService spawns one host-side proxy thread that drains its FIFO and posts IB work requests. With LL combine pushing ~1k put + 60 atomicAdd FIFO entries per iter, that single thread is the wall-clock bottleneck on cross-node runs. Split the channel set across kNumProxyServices=4 separate services so the host-side dispatch parallelism scales linearly. SemaphoreIds and MemoryIds are scoped to a ProxyService, so: - addMemory() is broadcast to every service in the same global order so a single MemoryId still identifies the memory everywhere. - Each (peer_rank, channel_idx) is assigned to one proxy_idx via round-robin; the resulting PortChannel is built on that proxy and inherits its FIFO. The kernel is unchanged: the flat handle array routes the right way automatically. No kernel-level changes, no tuning of QP count, no new env knobs.
… 1 on Blackwell) Override at runtime with MSCCLPP_EP_NUM_PROXIES. N=8 is the knee on H100+IB; N>=12 collapses from CPU oversubscription. Intra-node LL is unchanged.
Add dist.barrier() + dist.destroy_process_group() in a finally block so non-zero ranks don't poll the TCPStore after rank 0 (the store server) exits, which produced noisy 'recvValue failed / Connection was likely closed' stack traces from ProcessGroupNCCL's HeartbeatMonitor. Also pass device_id to init_process_group in the internode test to silence 'Guessing device ID based on global rank' warnings.
Aligns with NCCL-EP's ep_bench convention (BW computed from average time across ranks). Previously we reported only the max time and computed BW per-rank, which made our numbers more pessimistic than NCCL-EP's.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Instantiate low-latency dispatch and combine kernels for hidden size 6656 and expose the shape through the functional and benchmark entry points. Reuse syncNamedBarrier for scheduler named barriers instead of inline PTX. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use communicator-backed direct mappings, remove RDMA paths, and flatten the HT source layout. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 15f71a84-4219-4ae9-a87e-e5fab4205de6
Extend direct-fabric HT launch and runtime support to 16 ranks. Decouple TMA contributor staging from global rank discovery and restore adaptive warp geometry. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Support E4M3 payloads with FP32 scales per 32 elements, return global token-major expert IDs with num_experts sentinels, and simplify QuantConfig. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 59760d2a-e65f-44b6-b2e6-9c271b834d7e
Use linear UE8M0 scale bytes per 32 elements, global token-major expert IDs, optimized Blackwell conversion targets, and updated validation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 59760d2a-e65f-44b6-b2e6-9c271b834d7e
Make the invalid token expert sentinel configurable and preserve tiny nonzero MXFP8 blocks for UE8M0 code zero. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 59760d2a-e65f-44b6-b2e6-9c271b834d7e
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
#836) ## Summary Adds `run_ep_bench_python.py` — an in-process low-latency EP benchmark that drives **both** the mscclpp EP (`MoECommunicator`) and NVIDIA NCCL-EP Python APIs through one shared paired dispatch→combine loop, so the two backends are timed with byte-for-byte identical methodology (matching `mscclpp_ep_bench.cu` output). ## Changes - New `test/python/ep/run_ep_bench_python.py` (renamed from `ep_bench_unified.py`), MPI/mpi4py bootstrap, `--backend {mscclpp,nccl,both}`, optional in-process CUPTI kernel timing. - `run_ep_bench.py`: dropped the redundant Python `mscclpp` (`ep_bench_ll.py`) backend; deleted `ep_bench_ll.py` (superseded by the unified script). - Updated the mscclpp LL API usage for the current (IPC-capable) runtime (`num_rdma_bytes=0`; removed `get_low_latency_rdma_size_hint`). - Docstring: clarified mscclpp LL supports both CUDA-IPC (NVLink) and RDMA/IB, and added single-node and validated 2-node (shared MNNVL fabric) launch examples with placeholder IPs/iface. Both backends run at 1 and 2 nodes over the shared NVLink fabric. --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Binyang Li <binyli@microsoft.com>
) The kineto per-kernel attribution in _kineto_kernel_us matched the "dispatch"/"combine" substring against the full demangled kernel name. The rank-major combineKernel is templated on DispatchLayout (combineKernel<.., DispatchLayout::RANK_MAJOR>), so its name contains the substring "dispatch" and was wrongly summed into the dispatch bucket. This only surfaced under --cuda-graph, where the combine kernel also appears in the dispatch profiling pass, doubling the reported dispatch kernel time (e.g. 16->32 us at 1 node). Match on the function name with template arguments stripped (text before the first "<") so DispatchLayout / CombineMode template params no longer collide.
This PR adds rank-major output support to the MSCCL++ EP low-latency (LL) path, wiring it through the CUDA kernels/runtime, Python bindings, and the Python tests/benchmarks. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Contributor
There was a problem hiding this comment.
Pull request overview
This PR introduces the MSCCL++ Expert-Parallel (EP) extension (LL + HT backends) with a Python-facing API, adds a unified benchmarking harness for multiple EP implementations (mscclpp, NCCL-EP, DeepEP, FlashInfer), and extends PortChannel/proxy plumbing to support remote 64-bit atomicAdd, alongside build/CI/docs updates to enable and validate the new functionality.
Changes:
- Add
src/ext/epEP extension (C++/CUDA + nanobind) and Python frontend (python/mscclpp/ep) for MoE dispatch/combine. - Add EP benchmark backends + helpers under
test/python/ep/, including an optional CUPTI kernel timer. - Add PortChannel
atomicAddtriggers and connection implementations (CUDA-IPC / IB / Ethernet), plus new unit tests; update build options, docs, and CI defaults.
Reviewed changes
Copilot reviewed 74 out of 74 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
| test/python/ep/ep_bench_nccl.py | New NCCL-EP backend for unified EP benchmarking harness. |
| test/python/ep/ep_bench_mscclpp.py | New MSCCL++ EP backend benchmark driver with optional validation and CUDA graph options. |
| test/python/ep/ep_bench_flashinfer.py | New FlashInfer MoeAlltoAll backend adapter for benchmark comparisons. |
| test/python/ep/ep_bench_deepep.py | New DeepEP v2 ElasticBuffer backend adapter for benchmark comparisons. |
| test/python/ep/ep_bench_common.py | Shared MPI/torch bootstrap, input generation, dtype helpers, and validation utilities for EP benches. |
| test/python/ep/cupti_kernel_timer.cpp | Optional in-process CUPTI timer library for low-perturbation kernel timing. |
| test/python/ep/CMakeLists.txt | Standalone CMake build for EP bench binary + CUPTI timer. |
| test/mscclpp-test/CMakeLists.txt | Add tma_pipeline_perf test binary gated to CUDA and sm90+ arch lists. |
| test/mp_unit/port_channel_tests.cu | Add concurrent PortChannel atomicAdd test coverage across IPC/IB/Ethernet. |
| test/mp_unit/mp_unit_tests.hpp | Add testAtomicAdd declaration to PortChannelOneToOneTest. |
| src/ext/ep/runtime_base.hpp | New EP runtime base class defining common topology/availability surface. |
| src/ext/ep/README.md | New EP extension architecture/build/validation documentation. |
| src/ext/ep/moe_runtime.hpp | Declare createMoERuntime(...) factory for LL/HT EP backends. |
| src/ext/ep/moe_runtime.cc | Implement createMoERuntime(...) to construct LL or HT runtime. |
| src/ext/ep/low_latency/config.cuh | LL dispatch configuration, workspace views, kernel config caching utilities. |
| src/ext/ep/ll_runtime.hpp | LL runtime class declaration + dispatch/combine host entry points. |
| src/ext/ep/ll_runtime.cc | LL runtime implementation: symmetric buffer setup + dispatch/combine wrappers. |
| src/ext/ep/include/quantization.cuh | FP8 E4M3 quantization helpers for LL dispatch. |
| src/ext/ep/include/launch.cuh | Cooperative launch macros and rank-switch helpers for EP kernels. |
| src/ext/ep/include/exception.cuh | EP exception/assert/CUDA_CHECK helpers for torch-free kernel/runtime code. |
| src/ext/ep/include/constants.cuh | EP constants/timeouts/alignment + Torch restriction removal for kernels. |
| src/ext/ep/include/api.cuh | EP private kernel API: enums, workload/context structs, dispatch/combine declarations. |
| src/ext/ep/ht_runtime.hpp | HT runtime class declaration for direct-mapped pool + dispatch/combine protocol. |
| src/ext/ep/ht_runtime.cc | HT runtime implementation: pool mapping, notify/dispatch/combine plumbing. |
| src/ext/ep/high-throughput/runtime.cu | HT runtime helper kernels (e.g., barrier). |
| src/ext/ep/high-throughput/layout.cu | HT routing layout construction kernel (counts + token-in-rank). |
| src/ext/ep/high-throughput/dispatch.cu | HT notify/dispatch kernels with cooperative grid sync. |
| src/ext/ep/high-throughput/config.cuh | HT config constants and pool sizing helpers. |
| src/ext/ep/high-throughput/combine.cu | HT TMA pipeline combine kernel implementations. |
| src/ext/ep/config.hpp | Shared EP config helpers + LL symmetric buffer layout definitions. |
| src/ext/ep/CMakeLists.txt | Build mscclpp_ep_cpp nanobind extension; enforce sm90+ arch selection. |
| src/ext/ep/bindings.cpp | nanobind module exposing MoERuntime + LL/HT raw-pointer entry points. |
| src/ext/CMakeLists.txt | Wire EP extension into overall ext build when enabled. |
| src/core/unix_socket.cc | Add shutdown helper and destructor cleanup for UnixSocketServer lifecycle. |
| src/core/port_channel.cc | Add proxy handling for atomicAdd triggers; adjust stop/drain behavior. |
| src/core/include/unix_socket.hpp | Add UnixSocketServer destructor and shutdown declaration. |
| src/core/include/context.hpp | Add proxy atomic stream/context members and atomicAdd API to CudaIpcStream. |
| src/core/include/connection.hpp | Add atomicAdd virtual API to connection implementations. |
| src/core/include/communicator.hpp | Update comments related to last-recv-item lifecycle behavior. |
| src/core/context.cc | Implement CudaIpcStream destructor and document atomic stream sync caveats. |
| src/core/connection.cc | Implement Connection::atomicAdd and transport-specific atomicAdd paths. |
| src/core/communicator.cc | Refactor ordered recv/connect/semaphore futures without helper wrapper. |
| src/core/atomicadd_kernel.cu | Add device kernel + CudaIpcStream::atomicAdd implementation. |
| python/test/test_gpu_buffer_pool_nvls_zero.py | New NVLS zero-copy allreduce timing test using GpuBufferPool + CUDA graphs. |
| python/mscclpp/ep/utils.py | New EP Python utility helpers (pointer views, object broadcast/gather, etc.). |
| python/mscclpp/ep/types.py | New EP public Python dataclasses/types for dispatch/combine handles and config. |
| python/mscclpp/ep/communicator.py | New high-level Python MoECommunicator dispatch/combine API. |
| python/mscclpp/ep/_cpp.py | EP extension loader/shim exporting nanobind symbols. |
| python/mscclpp/ep/init.py | EP package public exports. |
| python/mscclpp/_core/comm.py | Fix torch-group UniqueId broadcast to handle variable-size payloads. |
| python/csrc/core_py.cpp | Release GIL around Bootstrap send/recv bindings for better concurrency. |
| pyproject.toml | Enable EP extension build in scikit-build CMake defines. |
| include/mscclpp/port_channel_device.hpp | Add PortChannel device-side atomicAdd trigger emission helpers. |
| include/mscclpp/gpu.hpp | Extend HIP CUDA-driver API type/function shims (CUcontext, cuCtx*). |
| include/mscclpp/gpu_data_types.hpp | Add bf16<->fp32 packed conversions and guard bf16 alias conflicts. |
| include/mscclpp/fifo_device.hpp | Reserve trigger type==0 for atomicAdd semantics. |
| include/mscclpp/core.hpp | Add public Connection::atomicAdd API declaration. |
| docs/quickstart.md | Document .[cuda12,ep] extra and include EP in “all extras” example. |
| CMakeLists.txt | Add MSCCLPP_BUILD_EXT_EP option to top-level build configuration. |
| .github/workflows/mscclpp-lang.yml | Adjust CI LD_LIBRARY_PATH to avoid CUDA compat lib mismatch. |
| .github/workflows/lint.yml | Move lint jobs to ubuntu-24.04 runners. |
| .github/agents/device-code-reviewer.agent.md | Add device-code-reviewer agent guidance document. |
| .devcontainer/devcontainer.json | Update devcontainer base image to cuda13.0 tag. |
Comment on lines
+61
to
+72
| // Note: proxyAtomicStream_ is NOT synced here. The atomicAdd kernels are fire-and-forget | ||
| // operations that complete asynchronously on the GPU. Syncing them here would deadlock | ||
| // because sync() is called from the proxy thread while the main thread may hold the | ||
| // device context via cudaStreamSynchronize() on the test kernel's stream. | ||
| // | ||
| // TODO(#796): As a side effect, `Connection::flush()` does not order/complete pending | ||
| // remote `atomicAdd` operations on the CUDA-IPC transport, so PortChannel flush no | ||
| // longer guarantees that a peer kernel sees the updated value. EP currently relies on | ||
| // higher-level signaling (PortChannel signal/wait, FIFO drain) for ordering, but a | ||
| // correct fix needs a deadlock-free way to drain `proxyAtomicStream_` here. Carried | ||
| // over from the DeepEP `chhwang/dev-atomic-add-cleanup` cherry-pick; revisit before | ||
| // this lands on `main`. |
Comment on lines
+21
to
+35
| #if !defined(MSCCLPP_DEVICE_HIP) | ||
| // On CUDA, the proxy thread cannot launch kernels or perform stream operations on the | ||
| // primary context without deadlocking with the main thread's cudaStreamSynchronize(). | ||
| // The CUDA runtime uses a per-context lock; the main thread holds it while waiting for | ||
| // the test kernel, and the proxy thread needs it to launch the atomicAdd kernel. | ||
| // A separate CUDA context avoids this contention. | ||
| // | ||
| // TODO(#796): `dst` is a CUDA-IPC mapping registered in the primary/runtime context, so | ||
| // launching this kernel from `proxyAtomicCtx_` is technically UB (device pointers are | ||
| // context-scoped). It works in practice on current drivers because the IPC handle aliases | ||
| // the same physical allocation, but a correct fix would either (a) avoid the separate | ||
| // context (e.g. break the deadlock differently) or (b) re-open the IPC mapping inside | ||
| // `proxyAtomicCtx_`. Carried over from the DeepEP `chhwang/dev-atomic-add-cleanup` | ||
| // cherry-pick; revisit before this lands on `main`. | ||
| if (!proxyAtomicCtx_) { |
Comment on lines
+761
to
+771
| if (isAtomicAdd && received && size == sizeof(int64_t)) { | ||
| // Atomic add: receive the value, read-modify-write on GPU memory | ||
| int64_t addValue; | ||
| recvSocket_->recvUntilEnd(&addValue, sizeof(int64_t), &closed); | ||
| received &= !closed; | ||
|
|
||
| if (received) | ||
| mscclpp::gpuMemcpy(ptr + (recvSize / sizeof(char)), recvBuffer_.data(), messageSize, cudaMemcpyHostToDevice); | ||
| recvSize += messageSize; | ||
| if (received) { | ||
| int64_t current; | ||
| mscclpp::gpuMemcpy(reinterpret_cast<char*>(¤t), ptr, sizeof(int64_t), cudaMemcpyDeviceToHost); | ||
| current += addValue; | ||
| mscclpp::gpuMemcpy(ptr, reinterpret_cast<char*>(¤t), sizeof(int64_t), cudaMemcpyHostToDevice); | ||
| } |
Comment on lines
+646
to
+650
| // Step 3: Block 0 signals remote that all adds are done, flushes, then waits for remote. | ||
| if (blockIdx.x == 0) { | ||
| portChan.signal(); | ||
| portChan.flush(); | ||
| portChan.wait(); |
Clarify the EP backend phases and reduce duplicated low-level machinery: - use MSCCL++ memory-channel semaphores for HT device barriers - rename HT routing-count/exchange launchers and document each phase - size LL workspace exactly instead of reserving a fixed 32 MiB - replace blanket cooperative-launch macros with scoped launch configs - remove unused PTX helpers and the obsolete EP constants header - make EP default-on for supported CUDA/Python builds and detect concrete native GPU architectures Also remove invalid EP install-extra docs, a non-importable GPUBufferPool standalone test, and the orphan tma_pipeline_perf CMake target. Restore Unix socket lifecycle code to origin/main and fix torch subgroup bootstrap source handling. Validated with 32-GPU LL rank-major and 16-GPU HT benchmarks, plus 2/4-GPU HT and LL BF16/FP8/direct-send correctness runs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Apply the repository clang-format 18 style to the cooperative combine launch macro so tools/lint.sh passes for both C++ and Python. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
| const size_t poolHeaderBytes = high_throughput::Config::recvPoolHeaderBytes(numRanks_); | ||
| EP_HOST_ASSERT(recvX == static_cast<uint8_t*>(recvPoolPtrs_[rank_]) + poolHeaderBytes); | ||
|
|
||
| const int hiddenInt4 = hidden * xElementSize / sizeof(int4); |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.