Reinforcement learning for LLM inference scheduling. DQN agent learns to balance throughput, TTFT, latency, and memory pressure vs FIFO/SJF/priority baselines.
-
Updated
Jul 10, 2026 - Python
Reinforcement learning for LLM inference scheduling. DQN agent learns to balance throughput, TTFT, latency, and memory pressure vs FIFO/SJF/priority baselines.
GPU-resident, differentiable, vectorized 2D multi-agent simulator built on NVIDIA Warp with zero-copy PyTorch interop.
Add a description, image, and links to the vectorized-environments topic page so that developers can more easily learn about it.
To associate your repository with the vectorized-environments topic, visit your repo's landing page and select "manage topics."