-
Notifications
You must be signed in to change notification settings - Fork 2.8k
Pull requests: NVIDIA/TensorRT-LLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[None][fix] Waive test_nemotron_h_breakable_prefill_cuda_graph[ray-tp1]
#19355
opened Sep 17, 2026 by
farazkh80
Collaborator
Loading…
3 tasks done
[None][fix] Convert the DFlash capture tap to the buffer dtype
ci: full pre-merge approved
#19354
opened Sep 17, 2026 by
brnguyen2
Collaborator
Loading…
1 task done
[None][fix] Refuse paged-context attention when no fused kernel exists
#19353
opened Sep 17, 2026 by
brnguyen2
Collaborator
Loading…
1 task done
[None][fix] Warn when DFlash is used with disaggregated serving
api-compatible
Accepted LLM API contract change that is backwards-compatible
#19352
opened Sep 17, 2026 by
brnguyen2
Collaborator
Loading…
1 task done
[None][fix] Validate unused chat template controls and normalize router tool arguments
#19351
opened Sep 17, 2026 by
brnguyen2
Collaborator
Loading…
1 task done
[None][fix] Share speculative capture buffers across CUDA graph buckets
ci: full pre-merge approved
#19350
opened Sep 17, 2026 by
brnguyen2
Collaborator
Loading…
1 task done
[None][test] perf-sanity: trim DeepSeek-R1 and Nemotron-Ultra-V3 cases, add gen_only_no_context coverage
#19349
opened Sep 17, 2026 by
chenfeiz0326
Collaborator
Loading…
4 tasks done
[https://nvbugs/6737127][fix] Exchange the handle as its raw
CUDA_IPC_HANDLE_SIZE struct bytes via…
#19348
opened Sep 17, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[None][test] Add coverage for scaled_mm
#19347
opened Sep 17, 2026 by
StanleySun639
Collaborator
•
Draft
1 task
[None][fix] suppress GCP host detection xtrace in CI images
#19345
opened Sep 17, 2026 by
hanjingtian
Contributor
Loading…
[None][fix] Use bounded NVFP4 serving configs for MiniMax-M3 perf tests
#19344
opened Sep 17, 2026 by
yufeiwu-nv
Collaborator
Loading…
1 task done
[TRTLLM-16217][test] Add Rubin single-node serve perf cases
#19341
opened Sep 17, 2026 by
ruodil
Collaborator
Loading…
[None][test] Add nemotron_3.5_lightning_30b_nvfp4 and nemotron_3.5_lightning_30b_bf16 func and perf cases on Spark
#19340
opened Sep 17, 2026 by
JennyLiu-nv
Collaborator
Loading…
1 task done
[None][fix] Reconfigure cmake when build_wheel.py arguments change
#19339
opened Sep 17, 2026 by
brnguyen2
Collaborator
Loading…
[None][feat] Upgrade Blackwell cuteDSL MLA kernel for packed q heads and compact input
#19338
opened Sep 17, 2026 by
pengbowang-nv
Collaborator
Loading…
1 task done
[None][fix] Track KV cache locality in slot metadata
api-compatible
Accepted LLM API contract change that is backwards-compatible
[https://nvbugs/6777501][fix] Fix nemotron breakable cuda graph test parity check
#19335
opened Sep 17, 2026 by
dominicshanshan
Collaborator
Loading…
1 task done
[None][infra] Waive 1 failed cases for main in post-merge 2965
#19334
opened Sep 17, 2026 by
trtllm-agent
Collaborator
Loading…
[None][test] Add coverage for DeepseekV32ForCausalLM
#19332
opened Sep 17, 2026 by
StanleySun639
Collaborator
•
Draft
1 task
[None][feat] Support breakable CUDA graph for Qwen3.8 Flash-Next
#19330
opened Sep 17, 2026 by
Wanli-Jiang
Collaborator
Loading…
[TRTLLMINF-420][fix] Add Jenkins instance name to SLURM job
#19327
opened Sep 17, 2026 by
lyxxn0414
Loading…
1 task done
[TRTLLM-14024][feat] Prune CuTe DSL GEMM autotuner tactics with nvMatmulHeuristics
api-compatible
Accepted LLM API contract change that is backwards-compatible
#19326
opened Sep 17, 2026 by
peaceh-nv
Collaborator
Loading…
1 task done
[#18407][fix] Release PEFT adapter ownership on KVCacheV2Scheduler suspend, restore on resume
#19325
opened Sep 17, 2026 by
pujitha24
Loading…
1 task done
[None][test] Pair gen_only with gen_only_no_context as a control group in QA multinode perf list
ci: full pre-merge approved
#19324
opened Sep 17, 2026 by
fredricz-20070104
Collaborator
Loading…
1 task done
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.