Open-weight inference for coding agents.
Kourier.sh serves DeepSeek V4 Flash behind a single OpenAI-compatible endpoint, with Qwen 3.6 35B coming soon. Flat monthly plans — no per-token metering, no GPU fleet to manage.
Why we built this
The open-weight landscape gives us extraordinary models. Running them is the hard part. Capacity planning, routing, failover, prefix caching, queuing — most teams just want to call a model and get an answer back, not operate a GPU fleet.
We built Kourier.sh because agents are becoming the primary users of inference endpoints. Agents do not browse dashboards or compare providers. They call an endpoint, stream a response, and move on. The endpoint needs to be fast, reliable, and priced so agents can iterate without a token counter ticking toward a surprise bill.
Kourier.sh is that endpoint. Not a framework. Not a platform. A managed service that runs open models behind the interface every tool already speaks — the OpenAI Chat Completions API — so you spend your time building, not wiring.
Open models deserve good infrastructure
The weights are open. The serving should not be a moat. Running open models should be as simple as calling any proprietary API.
Agents are the primary user
Not a power-user feature. Not a developer-experience nicety. The architectural center of gravity for everything we build.
Pricing should not be a puzzle
Per-token metering penalizes exploration. Subscriptions let agents iterate, retry, and build without a counter ticking toward a surprise bill.
How we build
Six principles. None are aspirational — every one of them is load-bearing in the codebase.
Who's behind this
Kourier.sh is built and operated by Kourier AI Inc, a Texas corporation focused on open-model inference for coding.
Based across Europe and North America. Async-first. We ship from a single GitHub repo, and every architectural decision is documented. We'd rather show our work than sell you on it.
Not a platform company. Not chasing a billion-dollar exit. Building the best open-model inference endpoint for the developers and agents who use it — and getting back to work.
Where we're going
What shipped, what's in flight, what's planned. Quarterly updates, no surprises.
- 2026 H1Shipped
Initial public alpha
OpenAI-compatible endpoint live with DeepSeek V4 Flash. Marketing site, Cloud dashboard, and documentation shipped.
- 2026 H2In progress
Scale and refine
Add Qwen 3.6 35B, grow the GPU fleet, tune routing and rate limits against real agent workloads, and iterate on plans based on what users actually need.
Stay close
We ship in public. Read the code, follow the work, or just say hi.