Page 1 of 2

Singularity API — Reserved GPU Slots for DeepSeek V4 Flash

We host DeepSeek-V4-Flash-0731 — full weights, full 1M context — on dedicated GPUs. $0.49/hour reserves you a private lane: guaranteed 160 tok/s (typically 200–340). The guaranteed base — 200k+ output and 5M fresh input tokens per hour — is just the floor; actual throughput typically runs 2–4x that. Unlimited cached input at 98%+ measured hit rates. Max 32 users per node, ever. 3 minutes — honest answers help us size capacity.

Email

What would you mainly use it for?

What do you use today for DeepSeek-class models?

Your timezone

Which hours are you typically active (your local time)?

How many hours per day do you actively code / run agents?

A
B
C
D
E
F

How many days per week?

A
B
C
D

Estimated daily token consumption

A
B
C
D
E
F
Cline/Cursor/OpenRouter dashboards show this — a rough guess is fine

How many agents/requests do you typically run in parallel?

A
B
C
D

Typical context size per request?

A
B
C
D
E

Do you know your current cache hit rate?

A
B
C
D

Is $0.49 per slot-hour (guaranteed 160+ tok/s, specs above) attractive for your workload?

A
B
C
D
E

How many slot-hours would you realistically buy per month?

A
B
C
D
E

What matters most to you?

What would stop you from switching to us?

Want early beta access (free test hours in exchange for feedback)?

A
B