High-performance C++17 KV cache compression engine for LLM inference with Python bindings, SIMD acceleration, and multi-backend support.
python compression inference pytorch simd attention avx2 quantization kv-cache large-language-models llm cpu-optimization hadamard-transform machine-learning-performance
-
Updated
Jul 27, 2026 - Makefile