Support dsv4 model - #1593
Open
WANDY666 wants to merge 221 commits into
Open
Support dsv4 model#1593WANDY666 wants to merge 221 commits into
WANDY666 wants to merge 221 commits into
Conversation
Root cause of the historical cudagraph accuracy drop (gsm8k 0.96 -> 0.74, coherent-but-runaway generations; same 0.75 the pre-v5 fullslot_decode experiments worked around): _capture_decode warms up via copy.copy(infer_state), which SHARES decode_att_state. FlashMLASchedMeta is lazily planned at the first kernel call and written back onto that shared state, so the warmup pass locks a schedule planned for the dummy batch (seq=2); the capture pass then binds those stale scheduler tensors and every replay runs real requests with a tile schedule planned for near-empty kv (systematically under-read attention). Fix: reset_sched_meta_for_capture() hook on the nsa decode att state, invoked in both capture paths after warmup, so planning happens INSIDE the captured region and re-plans on every replay from live tensors. Validation (tp4, H200, prompt cache on): batch-1 greedy decode is now character-identical to eager; per-layer probe shows embed+swa layers bitwise equal under replay, benign rounding-class deltas only in compress layers, argmax unchanged. gsm8k 100q/128: cold 0.960/111s, warm 0.960/23.3s 100% hits (eager: 0.95-0.97, cold 141s / warm 50s). Batch-1 decode 20.4ms/token vs 142ms eager. 41/41 unit tests green. Codex review GO (incl. overlap-path symmetry). launch.sh: drop --disable_cudagraph, derive PYTHONPATH from the script dir (hardcoded tree path made a worktree launch silently serve main-tree code). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Graph-sandwich prefill (graphs capture dense ops only; attention/compressor run eagerly between segments) was already in-tree; enabling it exposed that HOLD-pad rows read the racing HOLD slot, making their hiddens nondeterministic and perturbing real rows via MoE expert batching (ulp-level, amplified ~1.9x/layer). Zero the pad rows' attention output. Residual greedy-trajectory divergence vs eager equals the fp4 marlin MoE kernel's own run-to-run reduction-order noise (eager-vs-eager control: 0/4 match), accepted statistically: gsm8k 100q cold 0.980/115.5s warm 0.960/25.9s (eager-baseline parity); batch-1 TTFT 1.86x at 46 tokens.
Resolve request-manager packaging, lazy model registration, paged reservation, MTP, PD, and service lifecycle conflicts. Preserve V4 packed-cache code for the subsequent cache migration.
WANDY666
force-pushed
the
support_dsv4_model
branch
from
September 28, 2026 02:54
a4a4c05 to
364a184
Compare
Keep the selected DSV4 model stack in one squashed change, with excluded commits documented in revert.md. Preserve the MTP hidden input buffer owned by decode CUDA Graph capture while retaining eager-mode cleanup.
WANDY666
force-pushed
the
support_dsv4_model
branch
from
September 28, 2026 03:06
364a184 to
8157442
Compare
# Conflicts: # lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/deepgemm_impl.py # lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/marlin_impl.py # lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/triton_impl.py # lightllm/common/basemodel/triton_kernel/fused_moe/grouped_fused_moe_ep.py # lightllm/common/basemodel/triton_kernel/fused_moe/moe_silu_and_mul.py # lightllm/common/basemodel/triton_kernel/fused_moe/moe_silu_and_mul_mix_quant_ep.py # lightllm/common/req_manager/__init__.py # lightllm/common/state_cache_manager/__init__.py
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.