AI Latency Tracker

· 45 providers · 4 regions

As of July 28, 2026, the fastest-responding AI inference API by edge latency (time-to-first-byte) is fireworks at 19 ms p50, measured from Asia (Tokyo) (100% uptime, n=288). Rankings differ by region — see the table below.

Independent, provider-neutral latency and uptime of AI inference APIs (OpenAI, Anthropic, Google, Mistral, DeepSeek, xAI, OpenRouter and more), measured directly from multiple regions and updated automatically. How this is measured →

Which AI API is fastest right now?

RegionFastest AI APIp50 TTFBUptime
Asia (Tokyo)fireworks19 ms100%
Europe (Germany)nscale98 ms100%
South America (São Paulo)openrouter58 ms100%
US (Central)fireworks28 ms100%

Which AI APIs are fastest overall?

Composite score 0–100 (100 = as fast as the region leader), averaged across all 4 measured region(s). How it’s computed →

#ProviderSpeed Index
1fireworks80
2openrouter71
3google57
4meta-llama49
5baseten38
6sambanova37
7inference-net36
8cohere35
9mistral35
10nscale33
11cerebras33
12nebius32
13replicate31
14aleph-alpha31
15reka30

Full ranking — Europe (Germany)

Edge latency (time-to-first-byte), p50/p95, last 24h. Other regions: see the per-region pages below.

#Providerp50 TTFBp95UptimeSamples
1nscale98 ms202 ms100%290
2openrouter99 ms297 ms100%290
3fireworks100 ms301 ms100%290
4aleph-alpha100 ms204 ms100%290
5nebius100 ms235 ms100%290
6mistral100 ms298 ms100%290
7google101 ms300 ms100%290
8reka101 ms299 ms100%290
9meta-llama102 ms302 ms100%290
10inference-net102 ms704 ms100%290
11baichuan182 ms639 ms100%290
12baseten197 ms299 ms100%290
13perplexity199 ms303 ms100%290
14replicate200 ms403 ms100%290
15cohere200 ms303 ms100%290

How fast is each region?

How fast is a specific provider?

Which provider is faster than which?