Skip to content

[None][perf] PDL on the missing decode edges - #17454

Draft
brb-nv wants to merge 2 commits into
NVIDIA:feat/m3_with_msafrom
brb-nv:user/brb/m3-perf-decode-pdl-edges
Draft

[None][perf] PDL on the missing decode edges#17454
brb-nv wants to merge 2 commits into
NVIDIA:feat/m3_with_msafrom
brb-nv:user/brb/m3-perf-decode-pdl-edges

Conversation

@brb-nv

@brb-nv brb-nv commented Aug 9, 2026

Copy link
Copy Markdown
Collaborator

@coderabbitai summary

Description

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

brb-nv added 2 commits August 9, 2026 13:26
…d it

Median consumer-start minus producer-end on the main stream was -64 ns for the
QKV GEMM into the fused producer and for the index-score kernel into the block
selector, against roughly -400 ns on the edges that already overlap. Both
consumers now launch with programmatic stream serialization and wait
immediately before their first dependent load, so their index arithmetic runs
in the producer's tail. Worth about 0.35 us per edge across 57 sparse layers.

Behaviour is unchanged: cudaGridDependencySynchronize is a no-op when the
kernel was not launched with PDL, and the launches follow getEnvEnablePDL as
the rest of the codebase does.

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
The scorer read seq_lens before griddepcontrol.wait. num_blocks derived
from it bounds an unmasked block_table load and an unmasked score store,
so a stale length is an out-of-bounds write, and seq_lens is patched
every step by on_update_kv_lens. That was safe only because nothing
between that patch and this kernel writes it. Move the load past the
wait and leave the barrier init and descriptor prefetch ahead of it as
the work PDL overlaps.

Also drop the launch_dependents at kernel entry. It is not a visibility
bug, since a consumer's wait blocks until this grid completes, but the
only consumer is the block selector, which reduces over every block this
grid writes and so cannot start early. Now that the selector is
PDL-launched, triggering at entry only left it resident and spinning in
its own wait for this kernel's whole lifetime. The implicit trigger as
the CTAs exit still gives it the launch overlap.

Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant