Skip to content

Status

Every number here comes out of the gateway’s own request log. A model nobody has called says so, instead of showing a percentage built on nothing.

All systems operational

trailing 24 hours
Success rate
100.00%
0 upstream failures
Median latency
2532 ms
p50 across all models
p95 latency
3797 ms
slowest 5% start here
Requests
12
7 models with traffic

Last 60 days

14 of 15 days with traffic were clean

2026-07-15today

A pale bar is a day with no recorded traffic, not a day of downtime.

By model

ModelStatusSuccessp50p95Requests
GPT-6 Astra400K contextToo few calls3199 ms3199 ms2
GPT-5.6 Sol400K contextNo traffic0
GPT-5.6 Terra400K contextNo traffic0
GPT-5.5400K contextNo traffic0
Claude Fable 5200K contextToo few calls1
Claude Opus 51M contextNo traffic0
Claude Sonnet 51M contextNo traffic0
Gemini 3.8 Flash1M contextToo few calls2862 ms3797 ms4
Gemini 3.7 Flash1M contextNo traffic0
Gemini 3.1 Pro Preview1M contextNo traffic0
Gemini 3.1 Flash Lite1M contextNo traffic0
Gemini 3 Flash1M contextToo few calls1641 ms1641 ms2
Grok 4.6256K contextNo traffic0
Grok 4.5256K contextNo traffic0
MiniMax M3200K contextToo few calls3406 ms3406 ms1
MiniMax M2.7 HighSpeed200K contextNo traffic0
MiniMax M2.7200K contextToo few calls1367 ms1367 ms1
MiniMax M2.5 HighSpeed200K contextNo traffic0
MiniMax M2.5200K contextToo few calls2185 ms2185 ms1

How this is measured

  • A failure means the upstream provider returned 5xx. A 402 over quota or a 429 over your rate limit is the gateway working as designed, and is not counted against it.
  • Latency is measured end to end at the gateway, including the upstream call, over successful requests only.
  • Below five calls in the window there is no percentage, because at that sample size one bad request reads as a twenty percent failure rate.