Kourier / Data
Measured inference.
With context.
Real requests, grouped by the work they carry. See response timing and output speed alongside the samples behind them—not a best-case benchmark or a performance guarantee.
01 / Production measurements
All dates and times UTCWindow: 2 Oct 2026, 17:50–3 Oct 2026, 17:50 UTC. Snapshot: 3 Oct 2026, 17:50 UTC.
Measurements present: 2 Oct 2026, 19:51–3 Oct 2026, 17:40 UTC. A selected window does not imply continuous observations.
Last successful ingestion: 3 Oct 2026, 17:49 UTC. Snapshots refresh approximately every five minutes.
First output
Collecting
Exact · median · includes reasoning
n = 18 · p95 requires 100 samples
First answer text
Collecting
Exact · median · excludes reasoning-only output
n = 4 · p95 requires 100 samples
End-to-end output
55.4 tok/s
Median · generated tokens include reasoning
n = 3,694
3,694 qualifying measurements from 3,871 successful production usage records (95.4% coverage). Median measured input: 88,800 tokens. What counts?
Generated tokens, including reasoning, divided by full gateway request duration. Includes the wait before output begins.
By input size
Input includes instructions, conversation, and tool context—not just your latest message. K = 1,024 tokens.
| Input tokens | Samples | Median |
|---|---|---|
| <8K | 35 | 66.0 tok/s |
| 8–32K | 381 | 57.8 tok/s |
| 32–128K | 2,390 | 57.3 tok/s |
| 128K+ | 888 | 48.6 tok/s |
Over time
Hourly buckets in UTC. Blank slots mean missing or insufficient measurements, not zero latency or downtime. Boundary hours may be partial.
View hourly measurements and sample counts
| Hour (UTC) | Samples | Median |
|---|---|---|
| 2026-10-02 17:00 | 0 | Insufficient samples |
| 2026-10-02 18:00 | 0 | Insufficient samples |
| 2026-10-02 19:00 | 39 | 66.1 tok/s |
| 2026-10-02 20:00 | 0 | Insufficient samples |
| 2026-10-02 21:00 | 0 | Insufficient samples |
| 2026-10-02 22:00 | 47 | 74.0 tok/s |
| 2026-10-02 23:00 | 238 | 51.3 tok/s |
| 2026-10-03 00:00 | 251 | 64.9 tok/s |
| 2026-10-03 01:00 | 78 | 61.7 tok/s |
| 2026-10-03 02:00 | 330 | 61.7 tok/s |
| 2026-10-03 03:00 | 480 | 48.7 tok/s |
| 2026-10-03 04:00 | 458 | 50.7 tok/s |
| 2026-10-03 05:00 | 309 | 63.2 tok/s |
| 2026-10-03 06:00 | 528 | 48.5 tok/s |
| 2026-10-03 07:00 | 435 | 47.8 tok/s |
| 2026-10-03 08:00 | 197 | 64.4 tok/s |
| 2026-10-03 09:00 | 0 | Insufficient samples |
| 2026-10-03 10:00 | 1 | Insufficient samples |
| 2026-10-03 11:00 | 0 | Insufficient samples |
| 2026-10-03 12:00 | 0 | Insufficient samples |
| 2026-10-03 13:00 | 58 | 76.4 tok/s |
| 2026-10-03 14:00 | 101 | 65.8 tok/s |
| 2026-10-03 15:00 | 113 | 71.9 tok/s |
| 2026-10-03 16:00 | 29 | 68.6 tok/s |
| 2026-10-03 17:00 | 2 | Insufficient samples |
02 / Repeatable checks
Production requests vary in size, cache reuse, and reasoning. Our downloadable probe makes six serial streaming requests: an initial and repeated prompt at each of three input sizes.
It records client-observed first output, first answer text, total duration, token usage, and finish reason. Initial requests are not guaranteed cold-cache; repeated prompts are not guaranteed cache hits. Shared load and network distance are uncontrolled.
Probe requests identify themselves as synthetic and are excluded from these production aggregates after instrumentation. We do not publish six runs as a representative benchmark or a quality score.
Uses your own API key and subscription. Six requests, one at a time, up to 768 generated tokens per request; reasoning uses that allowance. No automatic retries.
# Set KOURIER_API_KEY in your environment
node performance-probe.mjs > results.jsonOutput contains measurements, not your key or response content. A length-limited answer is not a successfully completed coding task.
03 / Methodology & limitations
- Population, not adoption
- One observation per correlated usage record for this model. Requests—not users—are weighted equally. Operator traffic may be present, and a few accounts may dominate. Caller-identified synthetic traffic is excluded; unlabeled historical probes cannot be reliably separated. These are not customer counts.
- Successful streaming requests only
- Timing and speed require successful usage and gateway outcomes, no client abort, more than one generated token, a positive observed output window, and valid duration. Non-streaming, failed, cancelled, single-token, and uncorrelated requests are excluded. Coverage uses successful usage records as its denominator—not all incoming API attempts. This is not an error-rate or uptime report.
- Timing boundaries
- Exact timings use a monotonic clock from gateway arrival, including account queueing, to observed upstream output. First output includes reasoning and tool events; first answer requires answer text. Client network travel and rendering are excluded. Historical estimates subtract the output-stream window from total duration and include the post-output tail.
- Speed and token accounting
- End-to-end speed divides generated tokens by the entire gateway duration, including the initial wait. It includes reasoning tokens and is not decode-only throughput. Cache and reasoning counters are recorded only when reported, not inferred as zero. In this measured cohort, cache counters are available for 18 requests and reasoning counters for 16.
- Sample thresholds and missing hours
- Medians require 20 samples; p95 requires 100, independently for each metric and bucket. These publication thresholds are not statistical confidence guarantees. Missing results stay blank. p95 means 95% of measured timings were at or below that value. No interpolation or synthetic gap filling.
- What this does not establish
- Observed performance depends on request mix, cache reuse, reasoning, and shared load. It does not establish coding quality, fleet capacity, availability, an SLA, or the speed your next request will receive. See service status for monitoring information.