-
Notifications
You must be signed in to change notification settings - Fork 2.7k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 1 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
[Bug]: Blackwell CuTe DSL GVR top-k decode kernel corrupts output / crashes (barrier-divergence race) on all-zeros pre_idx and tie-heavy rows
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#18338 In NVIDIA/TensorRT-LLM;Fail explicitly for unsupported dtypes in Rubin MoE atomic_add_func
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#18337 In NVIDIA/TensorRT-LLM;Allocate fused MoE bias parameters in the locality-domain device context
Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#18336 In NVIDIA/TensorRT-LLM;Synchronize bulk-async reductions before sC reuse in Rubin fused MoE finalize
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#18335 In NVIDIA/TensorRT-LLM;Fix scaled_mm CuTe compile contract for blockscaled persistent GEMM
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#18334 In NVIDIA/TensorRT-LLM;Validate preferred and fallback cluster shapes for Rubin BF16 preferred-cluster GEMM
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#18333 In NVIDIA/TensorRT-LLM;[Bug] trtllm-bench hangs forever at shutdown when --iteration_log points into a directory that does not exist
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#18297 In NVIDIA/TensorRT-LLM;KV-cache-aware router never matches LoRA or salted requests: lora_id is dropped and cache_salt is hashed with a different algorithm than the engine
Disaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Status: Open.#18156 In NVIDIA/TensorRT-LLM;- Status: Open.#18153 In NVIDIA/TensorRT-LLM;
[RFC]: Router Hint initiated P2P KV Cache Transfer Between TRT-LLM Workers
Disaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Status: Open.#18151 In NVIDIA/TensorRT-LLM;[Bug]: V1 MAX_UTILIZATION paused_requests are never processed in the PP executor loop
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18115 In NVIDIA/TensorRT-LLM;[RFC] DFlash2 for Qwen3.8 on consumer Blackwell: integration boundary before we propose a PR
Speculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#18085 In NVIDIA/TensorRT-LLM;