Simple plans. Transparent usage.
Pick a monthly plan. Pro at $20/month for hands-on coding, Max at $50/month. Every model serves one OpenAI-compatible API.
Pricing questions.
Pro gives you 300 effective requests per rolling 5h window and up to 2 concurrent requests. Max removes the request cap entirely with unlimited tokens and up to 5 concurrent sessions. Both plans include the full open model lineup.
GLM 5.2 and Qwen 3.6 35B. Both serve the same OpenAI-compatible API, so switching models is one parameter change. We add new open models as they ship.
No. Pro and Max are flat monthly plans. No per-token metering, no seat fees, no surprise bills. Pro covers 300 effective requests per rolling 5h window, Max covers unlimited tokens.
An effective request is one request counted against your rolling window quota. On Pro you get 300 per rolling 5h window. Max has no per-request cap, so the window does not apply.
No. Prompts and completions are processed in memory and discarded after the response returns. We retain only request metadata (timestamp, token counts, model id) needed for billing. We never train on your data.
Typical time-to-first-token is 47ms p50 and 180ms p99 under production load. Throughput scales horizontally with demand. We publish live latency and throughput numbers on the performance page.