AI Latency Tracker
· 45 providers · 4 regions
As of September 25, 2026, the fastest-responding AI inference API by edge latency (time-to-first-byte) is fireworks at 19 ms p50, measured from Asia (Tokyo) (100% uptime, n=289). Rankings differ by region — see the table below.
Independent, provider-neutral latency and uptime of AI inference APIs (OpenAI, Anthropic, Google, Mistral, DeepSeek, xAI, OpenRouter and more), measured directly from multiple regions and updated automatically. How this is measured →
Is any AI API down right now?
2 confirmed outage(s) in the last 24 hours: doubao, Google
45 providers · probed every 5 minutes from 4 regions · last probe Sep 25, 09:13 UTC
None of the 10 providers with a machine-readable status page reports an open incident (checked Sep 25, 08:17 UTC).
all probes succeededsome probes failedhalf or more failedno data
Hour by hour for the last 72 hours, all regions combined; hover a cell for the count. Each provider page shows the same view per region. Full incident log →
Which AI API is fastest right now?
| Region | Fastest AI API | p50 TTFB | Uptime |
|---|---|---|---|
| Asia (Tokyo) | fireworks | 19 ms | 100% |
| Europe (Germany) | nscale | 98 ms | 100% |
| South America (São Paulo) | openrouter | 60 ms | 100% |
| US (Central) | fireworks | 31 ms | 100% |
Which AI APIs are fastest overall?
Composite score 0–100 (100 = as fast as the region leader), averaged across all 4 measured region(s). How it’s computed →
| # | Provider | Speed Index |
|---|---|---|
| 1 | fireworks | 81 |
| 2 | openrouter | 71 |
| 3 | 59 | |
| 4 | meta-llama | 46 |
| 5 | inference-net | 42 |
| 6 | baseten | 42 |
| 7 | cohere | 37 |
| 8 | sambanova | 36 |
| 9 | mistral | 36 |
| 10 | cerebras | 34 |
| 11 | nscale | 34 |
| 12 | nebius | 32 |
| 13 | aleph-alpha | 32 |
| 14 | replicate | 32 |
| 15 | reka | 31 |
Full ranking — Europe (Germany)
Edge latency (time-to-first-byte), p50/p95, last 24h. Other regions: see the per-region pages below.
| # | Provider | p50 TTFB | p95 | Uptime | Samples |
|---|---|---|---|---|---|
| 1 | nscale | 98 ms | 200 ms | 100% | 288 |
| 2 | fireworks | 98 ms | 202 ms | 100% | 288 |
| 3 | aleph-alpha | 99 ms | 293 ms | 100% | 288 |
| 4 | 99 ms | 203 ms | 100% | 288 | |
| 5 | openrouter | 99 ms | 202 ms | 100% | 288 |
| 6 | nebius | 100 ms | 201 ms | 100% | 288 |
| 7 | mistral | 100 ms | 205 ms | 100% | 288 |
| 8 | inference-net | 101 ms | 800 ms | 100% | 288 |
| 9 | meta-llama | 102 ms | 291 ms | 100% | 288 |
| 10 | baichuan | 191 ms | 296 ms | 99.7% | 288 |
| 11 | baseten | 198 ms | 301 ms | 100% | 288 |
| 12 | replicate | 198 ms | 398 ms | 100% | 288 |
| 13 | perplexity | 199 ms | 304 ms | 100% | 288 |
| 14 | glm | 199 ms | 301 ms | 100% | 288 |
| 15 | cerebras | 200 ms | 389 ms | 100% | 288 |
How fast is each region?
- Fastest AI API from Asia (Tokyo)
- Fastest AI API from Europe (Germany)
- Fastest AI API from South America (São Paulo)
- Fastest AI API from US (Central)
How fast is a specific provider?
- ai21 latency by region
- aleph-alpha latency by region
- anthropic latency by region
- baichuan latency by region
- baseten latency by region
- cerebras latency by region
- cohere latency by region
- deepinfra latency by region
- deepseek latency by region
- doubao latency by region
- ernie latency by region
- featherless latency by region
- fireworks latency by region
- friendli latency by region
- glm latency by region
- google latency by region
- groq latency by region
- hunyuan latency by region
- hyperbolic latency by region
- iflytek latency by region
- inference-net latency by region
- kimi latency by region
- meta-llama latency by region
- minimax latency by region
- mistral latency by region
- nebius latency by region
- novita latency by region
- nscale latency by region
- openai latency by region
- openrouter latency by region
- perplexity latency by region
- qwen latency by region
- reka latency by region
- replicate latency by region
- sambanova latency by region
- sarvam latency by region
- sensenova latency by region
- siliconflow latency by region
- stepfun latency by region
- targon latency by region
- together latency by region
- upstage latency by region
- writer latency by region
- xai latency by region
- yi-01ai latency by region