AgentDish directory

llm-serving

Accepted listings with this tag.

Listing Category Score Trend Checked
#1030 ↓ -6
LLM inference at scale

An open-source handbook for production LLM serving and inference at scale, covering GPU fundamentals, KV cache, batching, quantization, speculative decoding, and engines like vLLM, SGLang, and TensorRT-LLM.

Developer Tools / AI Infrastructure 84 ↓ -6 97 days ago Details

A research repo that replays real Claude Code and Mooncake traces to test prefix-cache eviction policies and compare them against LRU.

Developer Tools / Code Assistant 82 ↓ -2 43 hours ago Details

A detailed vLLM blog post explaining speculative decoding on AMD GPUs, with mechanics, supported drafting methods, configuration guidance, tuning advice, and benchmark discussion.

AI/ML / Inference Optimization 82 ↓ -2 4 days ago Details