We host DeepSeek-V4-Flash-0731 — full weights, full 1M context — on dedicated GPUs. $0.49/hour reserves you a private lane: guaranteed 160 tok/s (typically 200–340). The guaranteed base — 200k+ output and 5M fresh input tokens per hour — is just the floor; actual throughput typically runs 2–4x that. Unlimited cached input at 98%+ measured hit rates. Max 32 users per node, ever. 3 minutes — honest answers help us size capacity.
Cline/Cursor/OpenRouter dashboards show this — a rough guess is fine