LLM Leaderboard · Metacritic for Models

One page to pick the right model for the job: public quality evals as visible per-source bars, per-million-token pricing, refusal-rate data by domain, and sliders for what you care about. Move a slider, the ranking re-computes live.

Data as of loading – models tracked Refreshed daily

Quality: LMArena leaderboard (CC-BY-4.0). Pricing and context: OpenRouter public API (prices include OpenRouter's service fee, so they sit slightly above direct vendor prices). Refusal rates: moderationbias.com independent audit, LLM-judged, used as a labeled proxy. Param counts: Hugging Face Hub API.

Models with no refusal data are never excluded by this slider.
#ModelScoreArenaCoding MathPrice / 1MContextDeploy Why

Methodology

The bars

Refusal data

Refusal rates come from an independent moderation audit that prompts each model identically at temperature 0 and has an LLM judge classify refusals. It is a moderation-bias study, not a safety benchmark, so we present it as a labeled proxy with its source linked, and we never filter out models that lack coverage.

VRAM estimates

For open-weight models with verified parameter counts: VRAM = params x bytes-per-param x 1.15 (FP16 = 2 bytes, INT8 = 1, INT4 = 0.5; the 1.15 covers KV cache and runtime overhead). Rough estimate, not a deployment guarantee.

Honest caveats