llm-router live status

models →◐ ~5s fetch
valkey ready 1ms·config_rev —·inflight 4·generated 2026-10-11 06:39 UTC

Request history

completed requests per backend, from the spend log
dim·backendsusers·bucket·period·render
loading request history…

Request history: completed requests per series per bucket. Buckets are naive-UTC (a "day" starts at 00:00 UTC, matching the budget counter) and the last bucket is the current one — still filling in. 30d always uses day buckets (720 hourly buckets would exceed the cap). Proxy reads the router spend log (router era only, data from 2026-08-20); user reads ClickHouse llm_router.logs (full history incl. the litellm backfill) and charts the top 10 users by requests, folding the rest into others; user series are named from the users table (alias or email) when it knows them. The user selector works in both dimensions: in user it picks the series, in proxy it filters every per-proxy series down to the rows those users produced (an active filter shows as a chip; click it to clear). Hover for the per-bucket readout; click a line, bar or legend chip to isolate that series. The chart draws at most 12 series at once (the palette limit, not a data limit). Polled with the page; the server keeps each answer for 60s.

8/10degraded

GLM

inflight 41 unreachable · 1 locked
healthyglm-prx1glm-prx1.ismart.org0 · 0.58/s · 1peak—
healthyglm-prx2glm-prx2.ismart.org0 · 0.42/s · 1peak—
lockedglm-prx3glm-prx3.ismart.org0 · 0/s · 0peak—
zai 1309 plan expired — recharge required
locked since 2026-10-08 16:51 UTC
healthyglm-prx4glm-prx4.ismart.org0 · 0.58/s · 3peak—
healthyglm-prx5glm-prx5.ismart.org0 · 1.3/s · 3peak—
healthyglm-prx6glm-prx6.ismart.org0 · 0.62/s · 2peak—
healthyglm-prx7glm-prx7.ismart.org0 · 0.18/s · 1peak—
unreachableglm-prx8glm-prx8.ismart.org0 · 0/s · 0peak—
3 consecutive upstream connect failures (2026-10-11T05:41:29.317Z) — rechecking every 5m
since 2026-10-11 05:47 UTC
healthyglm-prx9glm-prx9.ismart.org0 · 0/s · 0peak—
healthyglm-prx10glm-prx10.ismart.org0 · 1.1/s · 3peak—
2/2healthy

Anthropic

inflight 0
healthycld-prx1cld-prx1.ismart.org0 · 0/s · 0peak—
healthycld-prx2cld-prx2.ismart.org0 · 0/s · 0peak—
1/1healthy

Kimi

inflight 0
healthykimi-prx1kimi-prx1.ismart.org0 · 0/s · 0peak—
1/1healthy

MiniMax

inflight 0
healthyminimax1api.minimax.io/anthropic0 · 0/s · 0peak—
0/1down

OpenCode

inflight 01 throttled
throttledopencode-prx1opencode-prx1.ismart.org0 · 0/s · 0peak—
1308 5h usage window exhausted
since 2026-10-08 17:04 UTC
earliest reset:

QPS column: now = queries in the current second · avg60 = queries/sec averaged over the last 60s (e.g. 0.47/s ≈ 28 queries/min) · peak60 = queries in the single busiest second of the last minute (what the ZAI per-second concurrency cap charges).

CREDITS 5h / 7d column: settled credits per backend against its 24,000-credits/5h and 120,000-credits/7d caps (each 85% of ZAI's own allowance) — read live from the SAME hourly buckets the budget settles into, so the display and the admission gate cannot disagree (a number here is at most ~30 s old). Amber when EITHER window is at 85% of its cap — spread the fleet before the cap sheds; red at 100% — already over, the backend is quarantining on its next admission check. SHED = our own credit budget is excluding this backend (until the instant in the tooltip); ZAI pause = ZAI itself is refusing the account (1310 quota / 1308 5h window / 1113-1313 lock) — a different action. — = the backend has no credit budget (request mode); its request counts stay on /health (volume) and in the chart above. stale = the credit counter could not be read; the numbers shown are the last good ones.