-
Notifications
You must be signed in to change notification settings - Fork 309
Pull requests: FlashML-org/FreeToken
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat(kvcache): 8-bit DSV4 window/compressed KV behind --kv-cache-dtype
#113
opened Aug 23, 2026 by
gdevenyi
Loading…
feat(moe): per-layer host-bank residency
#112
opened Aug 23, 2026 by
jason-fxz
Collaborator
Loading…
feat(server): expose swa_full_tokens_ratio as a startup CLI flag
#109
opened Aug 23, 2026 by
mkornreich
Loading…
feat(server): Add OpenAI /v1/models/{model_id} endpoint
#108
opened Aug 23, 2026 by
phillipmunn
Loading…
fix(dsv4): honor an explicit --max-prefill-length instead of silently forcing single-pass prefill
#105
opened Aug 23, 2026 by
avlp12
Loading…
feat(kvcache): 8-bit KV cache (q8_0 / fp8_e4m3) behind --kv-cache-dtype
#103
opened Aug 23, 2026 by
lucaspirola
Loading…
feat(laguna): native GGUF support for poolside Laguna (S/XS)
#102
opened Aug 23, 2026 by
lucaspirola
Loading…
perf(dsv4): compute the hyper-connection pre-norm in one kernel
#101
opened Aug 23, 2026 by
gdevenyi
Loading…
docs(install): document NCCL dev libraries for multi-GPU tensor parallelism
#90
opened Aug 23, 2026 by
alogotron
Loading…
perf(dsfp4): pick the grouped-prefill tile from route density on sm_89
#89
opened Aug 23, 2026 by
gdevenyi
Loading…
Fix: minor issue of tqdm lock leak during termination with TP
#88
opened Aug 23, 2026 by
xinze-zheng
Loading…
fix(fp8): round onto the e4m3 grid before the native float8e4nv downcast
#85
opened Aug 23, 2026 by
gdevenyi
Loading…
bench(bw): sweep decode batch size to measure cross-token expert dedup
#81
opened Aug 23, 2026 by
gdevenyi
Loading…
[3/3] perf(dsv4): add hardware-aware adaptive verification
#71
opened Aug 23, 2026 by
calvarado2004
Loading…
[1/3] feat(dsv4): add tensor-parallel DeepSeek-V4 runtime
#70
opened Aug 23, 2026 by
calvarado2004
Loading…
[2/3] feat(dsv4): implement exact DSpark speculative decoding
#69
opened Aug 23, 2026 by
calvarado2004
Loading…
Apple Silicon Metal backend (ft serve on macOS arm64)
#65
opened Aug 23, 2026 by
jasonkneen
Loading…
gemma4/gguf: accept a scalar attention.head_count_kv
#64
opened Aug 22, 2026 by
Rhycecomley
Loading…
Add Gemma-4 E-series (E2B/E4B) support: PLE, KV-layer sharing, double-wide MLP, and a dense GGUF parser
#59
opened Aug 22, 2026 by
mkornreich
Loading…
kernel: clamp zero-size host registrations; WDDM lock-pool hint on pin failure
#56
opened Aug 22, 2026 by
Glucksberg
Loading…
perf(cpu-moe): opt-in AMX tile GEMM for the deduped bf16 pass 1
#52
opened Aug 22, 2026 by
gdevenyi
Loading…
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.