Latency
45%Lower p95 latency scores higher only when it comes from authenticated streaming inference runs.
Live scoreboard for inference economics
Scoring methodology
The leaderboard reads from `benchmark_runs` in Neon Postgres. An hourly collection route refreshes the latest provider rows, but the UI computes a composite score only from authenticated inference measurements. When provider keys are not configured, public catalog probes keep source freshness visible while latency and throughput remain marked pending.
Collection contract
The live dataset ships through /api/leaderboard, while the scheduled collector writes fresh rows through /api/collect.
Lower p95 latency scores higher only when it comes from authenticated streaming inference runs.
Provider price is normalized per 1M output tokens so routing policies compare like-for-like.
Higher generated-token throughput improves score only after real completion measurements are available.
Current limitations