Skip to content

[None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host#16791

Open
hyukn wants to merge 1 commit into
NVIDIA:mainfrom
hyukn:yukunh/prepare-inputs-gettokens-oL
Open

[None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host#16791
hyukn wants to merge 1 commit into
NVIDIA:mainfrom
hyukn:yukunh/prepare-inputs-gettokens-oL

Conversation

@hyukn

@hyukn hyukn commented Jul 23, 2026

Copy link
Copy Markdown
Collaborator

Summary

request.get_tokens(0) marshals the whole C++ VecTokens (the entire sequence) into a fresh Python list of boxed ints — O(seq_len) host work that grows linearly with ISL. Three call sites in _prepare_tp_inputs / cuda_graph_runner pay this only to read a length or a small slice, and re-pay it every iteration. In chunked prefill the context loop re-marshals the full prompt for every chunk → O(L²/chunk) per request over the prefill. Normal decode is already immune (it uses get_last_tokens(0), O(1)); the waste is in the prefill/context path and the MTP first-draft paths.

Changes

  • New nanobind binding LlmRequest.get_tokens_range(beam, begin, end) — returns only the [begin, end) window, so nanobind copies just (end-begin) tokens (O(chunk)) instead of the whole VecTokens. Bounds are clamped.
  • context loop (_prepare_tp_inputs): use get_tokens_range(0, begin, end) for the current chunk and get_num_tokens(0) (O(1)) for the prompt length, instead of get_tokens(0) + slice + len(...).
  • first_draft loop (MTP): same — get_num_tokens(0) for the length, get_tokens_range for the last original_max_draft_len+1 tokens.
  • cuda_graph_runner first-draft branch: len(get_tokens(0))get_num_tokens(0).

Output is bit-identicalget_tokens_range returns exactly the same subrange the old get_tokens(0)[begin:end] produced.

Perf (DSv4-Pro disagg CTX worker, c3120, non-overlap, matched A/B; CTX nsys, iters 400–500, steady N=80)

_prepare_inputs p50 p90 success
baseline (this branch's base) 27.53 ms 43.29 ms
this PR 15.10 ms (−45%) 20.43 ms (−53%) 100% (3000/3000)

Each site drops from O(L) to O(1)/O(chunk); a synthetic microbench (bs128) shows each site flat across ISL after the fix (baseline get_tokens ~18 ms/iter at ISL 50K → ~0.05 ms). The binding was rebuilt and the new get_tokens_range verified present on the C++ LlmRequest; the full disagg run served at 100% success (functional parity).

Relationship to #16734

Complementary — different files, different cost. #16734 removes the overlap-scheduler device-scalar cudaStreamSynchronize in the DSV4 ctx sparse metadata (_prepare_inputs 268 ms → 19.96 ms); this PR removes the O(seq_len) get_tokens(0) host marshalling that remains. Together they target the two independent costs in _prepare_inputs.

🤖 Generated with Claude Code

…ng on the host

`request.get_tokens(0)` marshals the whole C++ VecTokens (entire sequence) into a
fresh Python list of boxed ints -- O(seq_len) host work that grows linearly with
ISL. Three call sites in `_prepare_tp_inputs` / `cuda_graph_runner` pay this only
to read a length or a small slice, and re-pay it every iteration. In chunked
prefill the context loop re-marshals the full prompt for every chunk -> O(L^2/chunk)
per request over the prefill. Normal decode is already immune (get_last_tokens(0),
O(1)); the waste is in the prefill/context and MTP first-draft paths.

Changes:
- New nanobind binding `LlmRequest.get_tokens_range(beam, begin, end)` that copies
  only [begin, end) (O(chunk)) instead of the whole VecTokens.
- context loop and first_draft loop: use get_tokens_range for the chunk and
  get_num_tokens(0) (O(1)) for lengths, instead of get_tokens(0) + slice.
- cuda_graph_runner first-draft branch: len(get_tokens(0)) -> get_num_tokens(0).

Output is bit-identical (get_tokens_range returns the same subrange). Measured on a
DSv4-Pro disagg CTX worker (c3120, non-overlap, CTX nsys): _prepare_inputs p50
27.53 ms -> 15.10 ms (-45%), p90 43.29 -> 20.43 ms; 100% request success. Complements
NVIDIA#16734 (which removes the overlap device-scalar sync, 268 -> 19.96 ms); together they
target the two independent costs in _prepare_inputs.

Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@hyukn
hyukn force-pushed the yukunh/prepare-inputs-gettokens-oL branch from 460954a to 1916433 Compare July 23, 2026 14:16
@hyukn
hyukn marked this pull request as ready for review July 23, 2026 14:16
@hyukn
hyukn requested review from a team as code owners July 23, 2026 14:16
@hyukn

hyukn commented Jul 23, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61304 [ run ] triggered by Bot. Commit: 1916433 Link to invocation

@coderabbitai

coderabbitai Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Adds a bounded get_tokens_range request binding and updates executor code to use direct token counts or limited token windows for chunked prefill, first-draft requests, and CUDA graph sequence-length calculations.

Changes

Token Range Retrieval

Layer / File(s) Summary
Token range binding
cpp/tensorrt_llm/nanobind/batch_manager/bindings.cpp
Adds get_tokens_range(beam, begin, end), which clamps bounds and returns a half-open token slice.
Model engine token windows
tensorrt_llm/_torch/pyexecutor/model_engine.py
Uses bounded token retrieval for context chunks and first-draft suffixes, and obtains prompt length from the request token count.
CUDA graph sequence sizing
tensorrt_llm/_torch/pyexecutor/cuda_graph_runner.py
Uses get_num_tokens(0) for first-draft sequence-length calculation.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Suggested reviewers: cascade812

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main perf change to avoid full token marshalling in prepare_inputs.
Description check ✅ Passed The description covers the issue, solution, and performance impact, with only the Test Coverage section left sparse.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants