Skip to content

Support dsv4 model - #1593

Open
WANDY666 wants to merge 221 commits into
mainfrom
support_dsv4_model
Open

WANDY666 wants to merge 221 commits into
mainfrom
support_dsv4_model

Conversation

@WANDY666

Copy link
Copy Markdown
Contributor

No description provided.

WANDY666 and others added 30 commits June 3, 2026 09:20
Root cause of the historical cudagraph accuracy drop (gsm8k 0.96 -> 0.74,
coherent-but-runaway generations; same 0.75 the pre-v5 fullslot_decode
experiments worked around): _capture_decode warms up via copy.copy(infer_state),
which SHARES decode_att_state. FlashMLASchedMeta is lazily planned at the first
kernel call and written back onto that shared state, so the warmup pass locks a
schedule planned for the dummy batch (seq=2); the capture pass then binds those
stale scheduler tensors and every replay runs real requests with a tile schedule
planned for near-empty kv (systematically under-read attention).

Fix: reset_sched_meta_for_capture() hook on the nsa decode att state, invoked in
both capture paths after warmup, so planning happens INSIDE the captured region
and re-plans on every replay from live tensors.

Validation (tp4, H200, prompt cache on): batch-1 greedy decode is now
character-identical to eager; per-layer probe shows embed+swa layers bitwise
equal under replay, benign rounding-class deltas only in compress layers,
argmax unchanged. gsm8k 100q/128: cold 0.960/111s, warm 0.960/23.3s 100% hits
(eager: 0.95-0.97, cold 141s / warm 50s). Batch-1 decode 20.4ms/token vs 142ms
eager. 41/41 unit tests green. Codex review GO (incl. overlap-path symmetry).

launch.sh: drop --disable_cudagraph, derive PYTHONPATH from the script dir
(hardcoded tree path made a worktree launch silently serve main-tree code).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Graph-sandwich prefill (graphs capture dense ops only; attention/compressor
run eagerly between segments) was already in-tree; enabling it exposed that
HOLD-pad rows read the racing HOLD slot, making their hiddens nondeterministic
and perturbing real rows via MoE expert batching (ulp-level, amplified
~1.9x/layer). Zero the pad rows' attention output.

Residual greedy-trajectory divergence vs eager equals the fp4 marlin MoE
kernel's own run-to-run reduction-order noise (eager-vs-eager control: 0/4
match), accepted statistically: gsm8k 100q cold 0.980/115.5s warm 0.960/25.9s
(eager-baseline parity); batch-1 TTFT 1.86x at 46 tokens.
Keep the selected DSV4 model stack in one squashed change, with excluded commits documented in revert.md. Preserve the MTP hidden input buffer owned by decode CUDA Graph capture while retaining eager-mode cleanup.
# Conflicts:
#	lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/deepgemm_impl.py
#	lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/marlin_impl.py
#	lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/triton_impl.py
#	lightllm/common/basemodel/triton_kernel/fused_moe/grouped_fused_moe_ep.py
#	lightllm/common/basemodel/triton_kernel/fused_moe/moe_silu_and_mul.py
#	lightllm/common/basemodel/triton_kernel/fused_moe/moe_silu_and_mul_mix_quant_ep.py
#	lightllm/common/req_manager/__init__.py
#	lightllm/common/state_cache_manager/__init__.py
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants