Kourier uses Google Analytics to understand aggregate website usage. Analytics is optional and is not used for targeted advertising. Read our Privacy Policy.

Open-weight inference for coding agents.

Kourier.sh serves DeepSeek V4 Flash behind a single OpenAI-compatible endpoint, with Qwen 3.6 35B coming soon. Flat monthly plans — no per-token metering, no GPU fleet to manage.

Why we built this

The open-weight landscape gives us extraordinary models. Running them is the hard part. Capacity planning, routing, failover, prefix caching, queuing — most teams just want to call a model and get an answer back, not operate a GPU fleet.

We built Kourier.sh because agents are becoming the primary users of inference endpoints. Agents do not browse dashboards or compare providers. They call an endpoint, stream a response, and move on. The endpoint needs to be fast, reliable, and priced so agents can iterate without a token counter ticking toward a surprise bill.

Kourier.sh is that endpoint. Not a framework. Not a platform. A managed service that runs open models behind the interface every tool already speaks — the OpenAI Chat Completions API — so you spend your time building, not wiring.

Open models deserve good infrastructure

The weights are open. The serving should not be a moat. Running open models should be as simple as calling any proprietary API.

Agents are the primary user

Not a power-user feature. Not a developer-experience nicety. The architectural center of gravity for everything we build.

Pricing should not be a puzzle

Per-token metering penalizes exploration. Subscriptions let agents iterate, retry, and build without a counter ticking toward a surprise bill.

How we build

Six principles. None are aspirational — every one of them is load-bearing in the codebase.

01
Open weights, no black boxes
DeepSeek V4 Flash is available today, with Qwen 3.6 35B coming soon. You know exactly which model answers every request. No mystery routing, no silent model swaps.
02
One endpoint, every model
OpenAI-compatible by design. Switch models by changing one parameter. If your tool speaks OpenAI, it already speaks Kourier.sh.
03
Flat pricing, no meters
Pro at $50/month, Max at $100/month. No per-token anxiety. Your agents can explore and iterate without watching a counter tick up.
04
Built for agents first
Rate limits, concurrency slots, and streaming tuned for automated workloads. Agents are the primary user, not a power-user afterthought.
05
Alpha means honest
Models, pricing, and rate limits will evolve as we scale. Every change is documented in the changelog. No silent regressions.
06
Show the work
Architectural decisions are documented in the open. We would rather be transparent than polished.

Who's behind this

Kourier.sh is built and operated by Kourier AI Inc, a Texas corporation focused on open-model inference for coding.

Based across Europe and North America. Async-first. We ship from a single GitHub repo, and every architectural decision is documented. We'd rather show our work than sell you on it.

Not a platform company. Not chasing a billion-dollar exit. Building the best open-model inference endpoint for the developers and agents who use it — and getting back to work.

Where we're going

What shipped, what's in flight, what's planned. Quarterly updates, no surprises.

  1. 2026 H1Shipped

    Initial public alpha

    OpenAI-compatible endpoint live with DeepSeek V4 Flash. Marketing site, Cloud dashboard, and documentation shipped.

  2. 2026 H2In progress

    Scale and refine

    Add Qwen 3.6 35B, grow the GPU fleet, tune routing and rate limits against real agent workloads, and iterate on plans based on what users actually need.

Stay close

We ship in public. Read the code, follow the work, or just say hi.