Skip to content

fix(cuda): support Turing GPUs - #24

Open
lukasrakauskas wants to merge 1 commit into
FlashML-org:mainfrom
lukasrakauskas:fix/sm75-turing
Open

fix(cuda): support Turing GPUs#24
lukasrakauskas wants to merge 1 commit into
FlashML-org:mainfrom
lukasrakauskas:fix/sm75-turing

Conversation

@lukasrakauskas

@lukasrakauskas lukasrakauskas commented Aug 22, 2026

Copy link
Copy Markdown

Adds NVIDIA Turing (sm_75) support for RTX 20-series GPUs:

  • Disable optional kernels that require sm_80+.
  • Use Turing-safe attention tile sizes.
  • Add sm_75 to kernel-cache builds.
  • Add backend and attention tests.

Verified on RTX 2060 SUPER with the Qwen3.6 NVFP4 model. The API server starts and serves
requests successfully.

The full test suite ran with 1,313 passed, 25 skipped, and 39 GPU-specific failures caused by RTX 2060 SUPER (sm_75) limitations and unavailable optional backends.

@gdevenyi

Copy link
Copy Markdown

Triaged while assembling a merged deployment branch for a 2× RTX 6000 Ada / 2× Xeon Gold 6526Y Linux box serving DeepSeek-V4-Flash with offloaded experts, to benchmark the open PRs together.

Turing (sm_75). Out of scope here (sm_89), not merged. Skimmed for shared-code risk and saw none that would affect Ada.

Flagging only so the absence of a report from me is not read as a problem found — I merged and benchmarked #30, #48, #56, #69, #70, #71 and #81, and left this one out deliberately.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants