LLM model scaling — real-time decode flow · VRAM ↔ RAM ↔ SSD

tick 0
×1.0
setup
model
quant
context
GPU
GPUs
system RAM
SSD
mmap from SSD