Skip to content

Latest commit

 

History

History
2 lines (2 loc) · 224 Bytes

File metadata and controls

2 lines (2 loc) · 224 Bytes

GPUCache

A PB-scale, ultra-low latency distributed GPU cache for AI inference. Built with Rust, NVIDIA DOCA, RDMA, and BF-4 DPUs to bridge GPU HBM and NVMe storage, eliminating the recompute tax for large language models.