Limited beta · 64 seats

Full-weight DeepSeek V4 Flash. Your own lane for an hour.

No shared rate limits. No throttling mid-session. Bring your real project, point your coding setup at it, and push it as hard as you want — free, for one hour.

We log zero prompts. Ever.

What you get

A real key, a real lane, real work.

Nothing to install on our side of the fence. You get credentials and a dedicated slice of the node for the full hour.

API key + base URLDrop it into Claude Code, opencode, Cline, Cursor, or hit the API directly. OpenAI-compatible.
A guaranteed laneNot a shared queue. A reserved slice of the node — your throughput doesn't collapse when others are busy.
~60 minutes, all-outCode on your own project at full tilt. When the hour ends, a two-minute feedback form and you're done.
The model · Shared Reserved Inference

A private GPU lane, without the private-GPU price.

Most AI serving makes you pick one of two bad deals: rent a whole GPU (truly dedicated, but pricey and idle half the day), or share a pay-per-token API (cheaper, but a crowded pool with rate limits and prices that move). Shared Reserved Inference is the middle path — one powerful GPU is split into equal, guaranteed lanes, and a small group reserves it together for a block of time. You get your own dedicated slice, and you split the cost with the cohort.

Per-token API// DeepSeek, OpenRouter…
  • Shared pool — you compete for capacity
  • Rate limits & throttling under load
  • Prices can change or spike anytime
  • No commitment
Shared Reserved Inference// what you're testing
  • Your own guaranteed lane — no crowd
  • No rate limits, no mid-session throttling
  • Flat, predictable hourly price
  • Full-weight model, 1M context
~$0.20–0.40/hr

per lane, for V4 Flash 0731

Rent a whole GPU// dedicated cloud GPU
  • Fully dedicated to you
  • $12–30+/hr — you pay for all of it
  • Idle time is wasted money
  • You run the setup yourself

A dedicated inference lane for less than a coffee.

Under $1/hour for your own guaranteed slice of a full-weight model — around $0.20–0.40/hour for V4 Flash 0731. On our last live run that worked out to 1.7–3.4× the token value you'd get spending the same amount on DeepSeek directly. This beta run is free — the pricing here is simply how the model works once it's live.

See the full benchmark
Zero logging

We don't read your code.

We store zero prompts and zero completions. During the beta we record only aggregate telemetry — request counts, latency, throughput, tokens, and cache-hit rate — plus whatever feedback you choose to give. Nothing you type or generate is ever kept.

Your prompts
Model completions
File contents
Request count & latency
Throughput & token totals
Cache-hit rate
From our last run

Built for cache-heavy coding.

Numbers measured on our previous saturated test of the same model. Your mileage in a live coding session will vary — that's what this beta is for.

6,200 tok/sAggregate output on cache-heavy coding workloads
97%+Prefix cache-hit rate under real repeat context
1MContext window, full model
FP4+FP8Native full weights — not a quantized cut-down

// full-weight DeepSeek V4 Flash 0731 · not a distill, not a quant

How it works

Sign up, get a key, code.

01Pick your slotJoin the Discord or fill the 60-second form, and tell us which of the candidate windows works for you.
02Get your keyIf you're selected, we send your API key and base URL by DM about 30 minutes before the slot starts.
03Go hard for an hourPoint your tools at the endpoint and work on your real project. Quick feedback form at the end.

Candidate windows: Wed Aug 20 & Thu Aug 21, at 14:00 or 16:00 UTC. Pick whatever fits your timezone in the form — we'll lock the final slot to whenever the most people can show up.

Ready?

64 seats. One hour. Full-weight, full speed.

Seats are limited — join early to lock your window.