0.2.0 — DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is now the live model on kourier.sh as deepseek-v4.1-flash. Earlier model ids are retired; update your client's model string to keep working.
v0.2.0AddedChangedRemovedkourier.sh now serves DeepSeek V4.1 Flash, DeepSeek's newest Flash model, on our own GPUs. It replaces DeepSeek V4 Flash on every plan, at the same flat monthly price.
What's in 0.2.0
Added
- DeepSeek V4.1 Flash — live now under the API model string
deepseek-v4.1-flash, with reasoning and tool calling - Every API surface works with it — OpenAI Chat Completions, the OpenAI Responses API (Codex), and the Anthropic Messages API (Claude Code), with streaming and tool calls
- Priority routing on Max — when the fleet is at capacity, Max requests are served ahead of Pro
Changed
- Model string — the model is now
deepseek-v4.1-flash: lowercase, vendor-first, matching how other providers name models - Context length — 262,144 tokens (256K), counting the prompt and generated output together
Removed
DSV4-Flash-0731andDSV4.1-Flash— both earlier model strings are retired. Requests that use them are rejected with HTTP 400.
Migrating
Change the model string to deepseek-v4.1-flash. Your API key, base URL, and
plan stay the same — existing keys were updated automatically.
OpenAI SDKs and curl
curl https://api.kourier.sh/v1/chat/completions \
-H "Authorization: Bearer $KOURIER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Hello"}]}'
Codex CLI — in ~/.codex/config.toml:
[profiles.kourier]
model_provider = "kourier"
model = "deepseek-v4.1-flash"
Claude Code — set both model variables:
export ANTHROPIC_MODEL="deepseek-v4.1-flash"
export ANTHROPIC_SMALL_FAST_MODEL="deepseek-v4.1-flash"
Oh My Pi — update the model id in ~/.omp/agent/models.yml and any
kourier/... model roles to kourier/deepseek-v4.1-flash.
GET https://api.kourier.sh/v1/models returns the exact model string your key
is granted. Full setup guides are in the docs.
Serve open models through one API.
DeepSeek V4.1 Flash today, with Qwen 3.6 35B coming soon. Two simple plans, $50/month or $100/month. No per-token metering.