Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -507,6 +507,7 @@ python3 download_pdfs.py # The code is generated by Doubao AI
|2025.12|🔥[**Grail-V/PSE**] Non-bijunctive Attention Collapse via POWER8 vec_perm — 8.8x CPU Inference Speedup(@Elyan Labs)|[[zenodo]](https://doi.org/10.5281/zenodo.14862410)|[[ram-coffers]](https://github.com/Scottcjn/ram-coffers) ![](https://img.shields.io/github/stars/Scottcjn/ram-coffers.svg?style=social)|⭐️ |
|2025.12|🔥[**llama-cpp-power8**] POWER8 optimizations for llama.cpp: vec_perm non-bijunctive collapse, IBM MASS integration, dcbt resident prefetch. 8.8x speedup over stock(@Scottcjn)|[[github]](https://github.com/Scottcjn/llama-cpp-power8)|[[llama-cpp-power8]](https://github.com/Scottcjn/llama-cpp-power8) ![](https://img.shields.io/github/stars/Scottcjn/llama-cpp-power8.svg?style=social)|⭐️ |
|2025.12|🔥[**RAM Coffers**] NUMA-aware weight banking for LLM inference. Maps brain hemisphere cognitive functions to NUMA topology for intelligent routing and selective prefetch(@Scottcjn)|[[github]](https://github.com/Scottcjn/ram-coffers)|[[ram-coffers]](https://github.com/Scottcjn/ram-coffers) ![](https://img.shields.io/github/stars/Scottcjn/ram-coffers.svg?style=social)|⭐️ |
|2026.06|[**Project Zero**] Zero-dependency C99 engine running BitNet ternary and GGUF dense in one binary; LUT-based ternary kernel hits 36 tok/s on Xeon (1.83x bitnet.cpp), no Python/BLAS(@shifulegend)|[[github]](https://github.com/shifulegend/project-zero)|[[project-zero]](https://github.com/shifulegend/project-zero) ![](https://img.shields.io/github/stars/shifulegend/project-zero.svg?style=social)|⭐️ |

### 📖Non Transformer Architecture ([©️back👆🏻](#paperlist))
<div id="Non-Transformer-Architecture"></div>
Expand Down