AgentDish directory
llm-serving
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#1030
↓ -6
LLM inference at scale
An open-source handbook for production LLM serving and inference at scale, covering GPU fundamentals, KV cache, batching, quantization, speculative decoding, and engines like vLLM, SGLang, and TensorRT-LLM. |
Developer Tools / AI Infrastructure | 84 | ↓ -6 | 97 days ago | Details |
|
A research repo that replays real Claude Code and Mooncake traces to test prefix-cache eviction policies and compare them against LRU. |
Developer Tools / Code Assistant | 82 | ↓ -2 | 43 hours ago | Details |
|
#1277
↓ -2
Speculative Decoding in vLLM on AMD GPUs
A detailed vLLM blog post explaining speculative decoding on AMD GPUs, with mechanics, supported drafting methods, configuration guidance, tuning advice, and benchmark discussion. |
AI/ML / Inference Optimization | 82 | ↓ -2 | 4 days ago | Details |