[CuTeDSL] Make autotuning more user friendly & new docs - #3408
Open
kainzhong wants to merge 7 commits into
Open
Conversation
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
kainzhong
force-pushed
the
improve_autotune
branch
from
August 5, 2026 21:57
475a4ab to
a34eb0b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds a lot of things to the overall CuTeDSL autotuning process to make it easier to use and more convenient:
derived_paramsto autotuning API so it computes derived arguments that must be decided during compile time but rely on autotuned parameters (whose values are different per configuration tuned)testing.autotune_jitable to decorate a class. This is for people who prefer to implement kernels as classes that fix the compile time params in__init__method and pass runtime params to__call__, so now they can do thisprune_configs_byin their autotuning APIautotune_suiteAPI so now user can provide a list of representative shapes and find the best overall kernel candidates with the ranking rule: sort according to how many times a configuration is accepted across all inputs, where "accepted" mean for a input case the configuration is in the topaccept_percentilefastest ones