One page to pick the right model for the job: public quality evals as visible per-source bars, per-million-token pricing, refusal-rate data by domain, and sliders for what you care about. Move a slider, the ranking re-computes live.
Quality: LMArena leaderboard (CC-BY-4.0). Pricing and context: OpenRouter public API (prices include OpenRouter's service fee, so they sit slightly above direct vendor prices). Refusal rates: moderationbias.com independent audit, LLM-judged, used as a labeled proxy. Param counts: Hugging Face Hub API.
| # | Model | Score | Arena | Coding | Math | Price / 1M | Context | Deploy | Why |
|---|
w x quality + (1-w) x value,
where quality = clamp((elo - 1200) / 300) and
value = 1 - clamp((log10(out_price) - log10(0.05)) / (log10(200) - log10(0.05))).
The weight w is your Quality-vs-price slider. Models with no public
price are scored on quality alone and flagged "price unknown".Refusal rates come from an independent moderation audit that prompts each model identically at temperature 0 and has an LLM judge classify refusals. It is a moderation-bias study, not a safety benchmark, so we present it as a labeled proxy with its source linked, and we never filter out models that lack coverage.
For open-weight models with verified parameter counts:
VRAM = params x bytes-per-param x 1.15 (FP16 = 2 bytes, INT8 = 1,
INT4 = 0.5; the 1.15 covers KV cache and runtime overhead). Rough estimate, not
a deployment guarantee.