Single-GPU fit: parameters × bytes-per-weight × 1.2 overhead, against nameplate VRAM and live prices. Rough by design — KV cache scales with context, and multi-GPU sharding changes everything. A screening tool, not a capacity plan.
| GPU | VRAM GB | On-demand $/hr | Where | Spot from |
|---|