Will a Hugging Face LLM fit your GPU or DGX? Plan weights, KV cache or recurrent state, context, TP/DP, and measured calibration
python cli capacity-planning gpu-memory model-serving huggingface kv-cache llm long-context vllm sglang gb10 vram-calculator dgx-spark tensor-parallel inference-memory
-
Updated
Jul 22, 2026 - Python