Methodology
WM Score rates every model 0–100 on price, quality, context, capabilities, reliability and host redundancy — not intelligence alone. Current version v1.3.
How the WM Score is built
Each factor family is normalised across the live market, so a score answers "how does this model compare to everything else available right now?" rather than "how did it do on one benchmark?".
- Quality
- Independent benchmark standing across reasoning, coding and general tasks.
- Price
- Blended input/output cost per million tokens, relative to the whole market.
- Context
- Usable context window — how much you can actually feed the model.
- Capabilities
- Vision, tool calling, structured output and reasoning support.
- Reliability
- Observed uptime and throughput from live host probes.
- Host redundancy
- How many independent providers serve the model — your fallback depth.
Weights are proprietary and versioned; we publish the version, the factor families, and every input we observe, so any score can be sanity-checked against the model's own page. Scores recompute on each market sync, and a score move is only published once it clears 5 points.
What counts as a move
We look for before-and-after changes in runtime code or production configuration: a default model replacement, provider initialization change, fallback change, router adoption, or rollback.
A model name in a README, benchmark, test, example, lockfile, or supported-model registry is not adoption by itself.
Evidence and review
Every public move links to the source commit and stores a relevant diff excerpt. Commit and pull-request language may explain a reason, but WhatModel does not infer reasons that are not stated.
Events are generated deterministically, then reviewed by an administrator before publication.
Confidence bands
- VERY_HIGH
- 95–100
- HIGH
- 85–94
- MEDIUM
- 70–84
- LOW
- 50–69
- REJECT
- Below 50
Important limitations
Public GitHub repositories are a biased and incomplete sample. A code change may not be deployed, and a repository may contain multiple applications or experimental paths.
Counts describe reviewed repository events. They should not be read as market share, production telemetry, or a recommendation that one provider is universally better.
